WritingBasetenBasetenpublished Jun 11, 2026seen 4w

Mercury 2 Is Now Available On Baseten

Open original ↗

Captured source

source ↗
published Jun 11, 2026seen 4wcaptured 4whttp 200method plain

Mercury 2, the first reasoning diffusion LLM, is now on Baseten Announcing our Series F . Learn more

AI models

Mercury 2, the first reasoning diffusion LLM, is now on Baseten

Inception's Mercury 2 is the fastest reasoning LLM, 5-10x faster than leading speed-optimized models at comparable quality.

Authors

Kumar Chellapilla

Lucas Bunzel

Bola Malek

Marylise Tauzia

Sid Sharma

Last updated June 11, 2026

Share

TL;DR Inception's Mercury 2 is now live on Baseten, making us the first inference platform to deliver production-grade diffusion LLMs (dLLMs) to developers. Mercury 2 runs over 1,000 tokens per second on widely-deployed NVIDIA GPUs, at less than half the cost, with comparable quality to Haiku and GPT-5 mini. It delivers real-time speed that used to require specialized hardware, with no custom chips and no lock-in. Customers like Augment Code are already running it in production, cutting costs by 90% and latency by 82% on critical workloads.

Traditional autoregressive LLMs generate tokens one at a time. Each token depends on the one before it, so generation is sequential by design, with a hard ceiling on speed. Over time, clever workarounds have been built, like speculative decoding and multi-head architectures, to predict several tokens at once. But these are inference-time patches on a model that's still autoregressive underneath. They ease the bottleneck without removing it.

Diffusion takes a different path. Our partner, Inception , is one of the leading labs building here, and instead of patching the constraint, dLLMs remove it. Rather than committing to one token at a time, it drafts the full output and refines it over several parallel passes, using the whole sequence to improve each part. Because this is built into how the model is trained and run, the speed is coming from the model itself, not a decoding optimization layered on top. It also opens a far richer design space, with more headroom for improvements ahead.

This isn't a marginal improvement. Augment Code, one of the first teams to run Mercury 2 in production on Baseten, cut costs by 90% and latency by 82% on a core part of their coding agent. And the gains aren't limited to coding: diffusion's speed opens up use cases that were previously very hard to serve, from real-time voice agents to sub-second tool routing and token-efficient subagents. "Our goal at Inception is to fundamentally redefine the economics and performance of LLMs so that they become more useful. Creating breakthrough architectures like dLLMs is only half the battle. Bringing them to market requires an equally innovative infrastructure partner. Baseten has built the gold standard for inference. Partnering with them to serve and optimize the model on NVIDIA hardware means our customers get the raw, parallel speed of Mercury 2 paired with the robust isolation, global scale, and compliance that enterprise production demands." Mercury 2 is Inception's flagship model and can generate over 1,000 tok/sec, speeds previously only possible with specialized AI inference chips. Why we partnered with Inception At its core, Baseten's job is to make it easy for customers to run inference efficiently at scale, regardless of the architecture they use. When we started to work with Inception, we quickly realized that there was a good fit. Mercury 2 is a model that enterprise teams are actively routing production traffic to for its speed, quality and cost effectiveness. Inception needed an inference partner that could handle enterprise-scale reliability, compliance requirements, observability, and customer isolation, and Baseten was a good match for that. "What excites me most about Inception is that they aren't just innovating on paper, they’ve built a high-performance architecture that enterprise teams are successfully routing production traffic to today. Diffusion LLMs present unique infrastructure challenges, and by combining Inception's breakthrough models with Baseten's enterprise-scale reliability, compliance, and isolation, we're making it seamless for developers to deploy these incredibly fast models into production." What the Baseten solution looks like Baseten is the infrastructure layer powering Inception’s Mercury API. Rather than building and operating their own inference platform, Inception routes customer traffic through Baseten, giving them enterprise-grade reliability and a growing set of platform capabilities without the operational burden. The deployment runs across NVIDIA GPUs including Hopper H100, Blackwell, etc. Because Mercury 2 delivers its speed at the model level, it doesn't depend on scarce specialized hardware; it runs on the widely available NVIDIA hardware enterprises already use, and it gets faster as those GPUs do. Baseten provisions always-on capacity with burst scaling support, which enables Inception to absorb traffic spikes without over-provisioning for steady-state load. Key platform capabilities Inception relies on include: Baseten Frontier Gateway for rate limiting per customer, request prioritization, and API routing

Metrics and observability

Autoscaling with configurable cron-based burst windows for peak traffic periods

Blackwell GPU cluster provisioned for voice and ultra-latency-sensitive workloads, targeting 150ms-250ms p50 end-to-end latencies Production results for Augment Code Augment Code is an AI-powered platform for enterprise software development, and they are one of the first teams running Mercury 2 in production on Baseten. Their use case is a good illustration of where diffusion LLMs shine. Augment's coding agents accumulate large context windows over the course of a session. When that context gets too large, the system needs to compress it so it summarizes the decisions made, files touched, issues unresolved, and next steps. This is called context compaction, and it's a hard problem to solve as it requires long-context understanding, fidelity, structured output generation, and low latency at the same time. The standard approach to address this challenge would be to use a frontier model for compaction, but it’s expensive and slow. That’s why Augment Code looked for alternatives and decided to try routing compaction to Mercury 2 as a dedicated subagent. "Building the best AI coding agent means using the right models for the right jobs. With Inception's Mercury 2 running on Baseten, we are able to take an...

Excerpt shown — open the source for the full document.

Notability

notability 6.0/10

Mercury 2 model release on Baseten platform.