WritingArcee AIArcee AIpublished Jun 18, 2025seen 4w

Deep Dive Afm 4 5b The First Arcee Foundational Model

Open original ↗

Captured source

source ↗

Arcee AI | Deep Dive: AFM-4.5B, the First Arcee Foundation Model

Trinity Large Thinking: Available on OpenRouter.

Try now ↗

ENTERPRISE

Research

COMPANY

Get API

Blog / Deep Dive: AFM-4.5B, the First Arcee Foundation Model

Deep Dive: AFM-4.5B, the First Arcee Foundation Model Mark McQuade ,

Lucas Atkins ,

Fernando Fernandes Neto ,

Charles Goddard ,

Varun Singh ,

June 18, 2025

Built for performance, compliance, and affordability.

Today marks a pivotal moment for Arcee AI and our customers: the launch of AFM-4.5B, the first Arcee Foundation Model. AFM-4.5B is the result of a deliberate, ambitious effort to deliver enterprise-grade AI that meets the real needs of today's organizations: performance, compliance, and affordability—at a scale and quality not previously available. For a quick taste, you can test AFM-4.5B in our playground and on Together.ai. Why We Built AFM-4.5B Our journey began in response to a pattern we saw across countless customer deployments. Over the years, we've helped organizations push AI further, driving better performance and lower costs through precision tuning and targeted post-training. But as AI adoption grew, we saw a set of inescapable pain points emerge. Performance and Size Gaps Edge-optimized models weren’t simply reliable enough for demanding tasks. Customers needed a model that could run on modest hardware, yet still deliver top-tier accuracy and robustness. Regulatory and Licensing Friction The most advanced models from major Chinese AI labs (Deepseek, Qwen, GLM, MiniCPM) offered impressive results, but rarely satisfied Western compliance standards, disqualifying them for regulated industries. Stagnant Western Alternatives Models from Meta (Llama) and Mistral, while solid, were quickly becoming outdated in relevance. The 3–10B parameter space was primarily served by models a year old or older, outpaced by newer research, data pipelines, and post-training strategies. Customers faced a hard choice: compromise on performance, compliance, or future flexibility. We knew there had to be a better way. The answer wasn’t a patchwork of tweaks or incremental improvements. We committed to a bold course: design and train a new model, from the ground up, for the world our customers actually operate in. The Making of AFM-4.5B: How We Trained It Training a foundation model of this scale is never simple. We took on the challenge not just to build a better model, but to prove that focus, rigorous data practices, and deep expertise could deliver a step-change in real-world utility. Uncompromising Data Quality

We knew that in order to build the strongest models possible, we needed the best possible training data. To achieve this, we partnered with DatologyAI , the leading experts in data curation, to assemble 6.58 trillion tokens of the most relevant, highest-quality data possible. Data curation for foundation models is hard. It's a frontier research problem—it's a comparatively new field, experiments are costly to run at scale, and small-scale results often aren't predictive of large-scale outcomes. It's also a frontier engineering problem—there's no established playbook for implementing a curation pipeline that can scale up to the trillions of tokens that needed to train competitive foundation models. We knew it just wouldn't make sense to try to tackle this ourselves. This is why we chose to partner with DatologyAI. DatologyAI's curation pipeline integrates a suite of proprietary algorithms—model-based quality filtering, embedding-based curation, target distribution-matching, source mixing, and synthetic data—and customizes them to generate a strong general-purpose dataset that also targets the capabilities we wanted our model to have. The results showed early: by 2 trillion tokens, AFM-4.5B was already outperforming competing models trained on dramatically larger, but noisier datasets.

Purpose-Built Infrastructure We utilized Amazon SageMaker Hyperpod and orchestrated training across 512 Nvidia H200 GPUs. This cloud infrastructure enabled us to experiment rapidly with various architectural variants, hyperparameter sweeps, and targeted interventions.

Expert Post-Training for Real-World Reliability AFM-4.5B’s clean foundation made it a prime candidate for our multi-stage post-training pipeline—built to adapt the model to enterprise demands through advanced fine-tuning, distillation, merging, and alignment techniques.

Who AFM-4.5B Is For AFM-4.5B is purpose-built for organizations that won't settle for compromises. ‍ Cost-Effective Inference Optimized for high throughput on CPUs, AFM-4.5B delivers GPU-tier results with efficient resource usage—ideal for both cloud and on-premise scenarios.

Flexible, Edge-to-Cloud Deployment From enterprise servers to mobile devices and IoT modules, AFM-4.5B scales effortlessly, supporting AI where you need it. On an Amazon EC2 c8g.8xlarge instance (Graviton4, 32 vCPUs), an 8-bit version of AFM-4.5 running on llama.cpp can deliver well over 100 tokens per second at batch size 4. A 4-bit version delivers over 200 tokens per second. This combination of high-quality generation and high CPU performance opens up cloud and edge use cases that were impossible until now. Post-Trained for Enterprise Tasks We designed AFM-4.5B with real-world use cases in mind and have refined it to handle a broad spectrum of workloads with precision and reliability.

Tools and Expertise, Ready for You We built AFM-4.5B to be more than just a foundation model; it's a launchpad for fast, reliable deployment in the real world. AgentHarness and Retrieval Toolkits Drop-in kits enable AFM-4.5B to power tool use, retrieval-augmented generation, and agentic reasoning—securely and on your infrastructure.

Rapid Domain Customization ‍ Fine-tune the model to your vertical in hours, not months. Our pipelines and documentation help you take full control without unnecessary overhead.

Deep Dive: The Post-Training Pipeline At the heart of AFM-4.5B’s real-world performance is its post-training stack—a layered strategy that surfaces and sharpens capabilities without sacrificing generality or introducing brittleness. It begins with midtraining, where we infused the model with high-leverage datasets (math, code, complex reasoning) and carefully selected samples from DatologyAI’s corpus. This step gave the model strong early instincts for precision and clarity. From there, we performed checkpoint...

Excerpt shown — open the source for the full document.

Notability

notability 6.0/10

First foundational model from Arcee, low HN traction.