WritingAI21 LabsAI21 Labspublished Apr 15, 2026seen 4w

Introducing Jamba Reasoning 3b

Open original ↗

Captured source

source ↗
published Apr 15, 2026seen 4wcaptured 4whttp 200method plain

Introducing Jamba Reasoning 3B: Tiny Model, Huge Possibilities | AI21

Skip to Main Menu

Skip to Main Content

Skip to Footer

Back to Blog

-->

Back to Blog

Today, we are introducing Jamba Reasoning 3B, a compact, open source reasoning model that redefines what is possible on-device—and marking the first in a series of new additions to the Jamba model family. Built on our novel SSM-Transformer architecture , with a context window length of 256K tokens and the ability to handle up to 1M tokens, Jamba Reasoning 3B introduces 2-5X efficiency gains over competitors such as DeepSeek, Google, Llama, and Microsoft, in addition to achieving leading intelligence benchmarks.

We are proud to release this model under the Apache 2.0 License , as part of our commitment to democratizing access to quality models; we are especially excited that, due to this model’s lightweight memory footprint, developers everywhere can download and run this technology right on their own devices, including iPhones, Androids, Macs, and PCs.

Try Jamba Reasoning 3B Now:

Hugging Face

Kaggle

LM Studio

Jamba Reasoning 3B’s release underscores NVIDIA’s recent proclamation that “small language models are the future of agentic AI. ” By successfully leveraging a KV cache that’s 8X smaller than the “vanilla” Transformer architecture, Jamba Reasoning 3B’s hybrid SSM-Transformer architecture keeps memory usage low, even as context grows. It can produce 40 tokens/second on an M3 MacBook Pro at context lengths of 32K tokens—making it a lean component for use within advanced agentic applications.

Model Fast Facts

License: Apache 2.0

Number of parameters: 3B

Context window length: 256K

Availability:

Download: 3B model, and quantized versions, available via Hugging Face and Kaggle

Local inference: 3B model, and quantized versions, available via LM Studio , and llama.cpp

A New Era of On-Device Intelligence

By packing leading intelligence scores, low latency, and long context into a compact reasoning model, Jamba Reasoning 3B opens up a new frontier for on-device experimentation and deployment—and this is just the beginning of what’s possible.

At AI21 Labs, we look forward to continuing to leverage our novel hybrid model architecture to release increasingly performant Jamba models, enabling enterprises to dial down the memory footprint on heavy workloads, without degrading quality. We believe these advancements hold the potential to build a more decentralized and democratic future, in which on-device computation enhances the economic viability of AI deployment across the whole ecosystem.

Tiny Models, Huge Possibilities

We look forward to how the community will use Jamba Reasoning 3B, from enterprise applications—such as enabling quick and local processing and entity extraction from legal or medical documents or equipping field technicians from a utility company with always-on access to manuals via their PCs—to personal on-device apps, such as productivity trackers and conversational and writing assistants securely tuned to your own file database.

Here are a few places where Jamba Reasoning 3B shines: Intelligence that doesn’t slow down : Due to its hybrid SSM-Transformer architecture, Jamba Reasoning 3B works more efficiently than pure Transformers models. While most Transformers-based significantly degrade in performance beyond context lengths of 32K tokens, Jamba Reasoning 3B handles far longer context lengths—including up to 1 million tokens—making it useful within advanced agentic AI systems or multi-modal applications where long context understanding is critical to output quality.

Leading intelligence : Jamba Reasoning 3B outperforms other on-device models from DeepSeek, Google, Meta, and Microsoft. It particularly shines on instruction-following tasks (IFBench) and general knowledge (MMLU-Pro and Humanity’s Last Exam), making Jamba Reasoning 3B both an efficient and intelligent model for use within advanced agentic workflows or on-device RAG applications. These results were achieved through a robust post-training pipeline, in which we applied a combined approach of alignment training techniques—such as RLVR, SFT, DPO, and GRPO—with our own proprietary methods, in order to ensure outstanding model quality.

Built for secure on-device use : Licensed under Apache 2.0, this model can be downloaded right to your computer or phone and customized on-device with your own files, for fully secure applications that keep running even if your internet doesn’t.

The Importance of On-Device Models for Enterprise AI

As enterprises integrate AI into operations, cloud-based large language models (LLMs) reveal economic inefficiencies, with high depreciation costs outpacing revenue. Research indicates 40-70% of AI tasks can be handled by small language models (SLMs) at 10-30x lower cost through intelligent routing. On-device SLMs like Jamba Reasoning 3B enable cost-effective, heterogeneous compute allocation—processing simple tasks locally while reserving cloud resources for complex reasoning. In agentic workflows, these SLMs can act as the on-device controller, orchestrating operations by activating cloud-based LLMs or external tools only as needed for optimal efficiency. This delivers low latency for real-time applications in manufacturing and healthcare, offline resilience for remote operations, and enhanced data privacy by keeping sensitive information on-device. This architecture ushers in a decentralized AI era, akin to the 1980s shift from mainframes to personal computers, empowering local computation while seamlessly integrating cloud capabilities for greater scalability. On-device models like Jamba Reasoning 3B are a key unlock for enterprise agentic applications, ensuring economic viability, reduced latency, and robust security.

Start Using Jamba Reasoning 3B

Available for download and local inference across Hugging Face , Kaggle , LM Studio , and llama.cpp , researchers and AI enthusiasts alike can enjoy experimenting with Jamba Reasoning 3B and pushing it to support new use cases—we can’t wait to see what you build!

Hugging Face

Kaggle

LM Studio

Discover more

Jun 25, 2026

Token spend isn’t going down. You need more than naive routing to manage it

Labs in Front

-->

Jun 24, 2026

Tipping the scales: Merging weak agents into a state-of-the-art deep researcher

Labs in Front

-->

Jun 4, 2026

First scale, then enrich: How the right execution strategy helped us reach state-of-the-art on...

Excerpt shown — open the source for the full document.

Notability

notability 7.0/10

New reasoning model release by AI21 Labs