Inference Engineering For Deepseek V4 Pro 0813
Captured source
source ↗Inference engineering for DeepSeek V4 Pro 0813
Model performance
Inference engineering for DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is a 1.7T-parameter open frontier model under the MIT license and is available for inference today on Baseten model APIs.
Authors
Model Performance Team
Last updated August 13, 2026
Share
DeepSeek V4 Pro 0813 weights were open-sourced today. This 1.7T-parameter frontier open model is MIT-licensed, meaning that every enterprise has the opportunity to run and fine-tune the model without restrictions, and is available today as a Model API on Baseten with zero data retention by default. ✕ DeepSeek AI published first-party quality benchmarks for DeepSeek V4 Pro 0813 Benchmarks were run using DeepSeek’s new coding harness, which was open-sourced alongside the new models as a developer preview. Alex Ker’s breakdown of the new harness reveals its plugin-first approach to operate at a higher level of abstraction than a standard coding loop. DeepSeek V4 Pro 0813 is a new open frontier model that sits alongside GLM-5.2 on quality on Artificial Analysis’ all-around intelligence index , though at a lower cost per task. This post summarizes the inference engineering work we did on DeepSeek V4 Pro 0813 to create a production-ready day zero API for the model. From preview to pro This release is an update of the DeepSeek V4 Pro Preview model, released in April, with additional post-training for code generation and agentic behavior. While the new model performs substantially better on both benchmarks and real-world use, those gains are achieved on the same base architecture. This is increasingly common – more and more gains in model quality are coming from post-training. The advantage of serving a new model with a familiar base architecture is that the inference setup can be heavily informed by work on the earlier model. For DeepSeek, we run the native MXFP4 weights on our proprietary inference engine within the Baseten Inference Stack using many of the features we originally built to support the preview model. However, it’s not quite as simple as swapping out the weights and shipping. First off, traffic patterns have evolved materially since April. The input and output sequence lengths are longer, cache hit rates are higher, and topics are more focused in today’s agentic coding landscape. This in turn affects the optimal configs across the inference stack, from parallelism to KV cache allocation to prefill-decode worker ratio. We configured every knob of the inference stack for an efficient balance of low latency and high throughput. Additionally, DeepSeek V4 Pro 0813 ships with some changes to the frontend and chat template to match the last four months of evolution in harness expectations and API specs. The documentation in the Hugging Face repository explains this change, which replaces a traditional Jinja-format chat template with a folder of scripts demonstrating input encoding and output parsing. Adherence to this spec is essential to high quality across tool calls and other structured outputs. Finally, the new model has slightly more parameters than the preview model as it ships with a DSpark speculator. While this doesn’t change the underlying architecture, it gives us the opportunity to support speculative decoding for higher TPS. While the native speculator is good for the most common use cases for the model – coding and agentic tasks – we would want to train a new speculator for most dedicated deployments on a narrower dataset that is representative of expected usage for that specific deployment. Production inference for DeepSeek V4 Pro 0813 DeepSeek V4 Pro 0813 is available today on Baseten Model APIs , as well as on dedicated deployments. DeepSeek V4 Pro 0813 brings frontier performance at lower prices than closed models: Input tokens (cache miss): $1.32/m
Input tokens (cache hit): $0.132/m
Output tokens: $3.96/m
You can get started today with a single API call: 1 import os 2 from openai import OpenAI 3 4 client = OpenAI( 5 api_key=os.environ[ "BASETEN_API_KEY" ], 6 base_url= "https://inference.baseten.co/v1" ) 7 8 response = client.chat.completions.create( 9 model= "deepseek-ai/DeepSeek-V4-Pro-0813" , 10 messages=[ 11 { "role" : "system" , "content" : "You are a helpful assistant" }, 12 { "role" : "user" , "content" : "Write Hello World in Python" }, 13 ], 14 stream= True , 15 reasoning_effort= "low" , 16 extra_body={ "thinking" : { "type" : "enabled" }} 17 ) 18 19 print (response.choices[ 0 ].message.content)
Subscribe to our newsletter Stay up to date on model performance, inference infrastructure, and more.
Explore Baseten today Start deploying Talk to an engineer
Related posts View all Model performance
Model performance Making Kimi K3 tokenization 18x faster for million-token agentic workloads
Michael Feil
Model performance How to build a day-0 API for Kimi K3
Model Performance Team
Model performance How we built the new fastest API for GLM-5.2
Alex Korte 6 others
Popular models Kimi K3
GLM-5.2 Fast
DeepSeek-V4-Flash-0731
Inkling
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Popular models Kimi K3
GLM-5.2 Fast
DeepSeek-V4-Flash-0731
Inkling
Whisper Large V3
NVIDIA Nemotron 3 Ultra
Explore all
Notability
notability 5.0/10Inference engineering guide for DeepSeek V4 Pro