WritingCoreWeaveCoreWeavepublished Sep 10, 2026seen Jul 29

Why AI Factories Need Proof Before Production

Open original ↗

Captured source

source ↗
published Sep 10, 2026seen Jul 29captured Jul 29http 200method plain

Co-designed AI factories for production AI | CoreWeave Blog

Announcement

Webinar

Podcast

GTC 2026

CoreWeave recognized as a Visionary in the Gartner® Magic Quadrant™ for Cloud AI Infrastructure. Read the report

Products

Data and storage

Infrastructure control

Runtime acceleration

Model and agent development

Mission control

Solutions

Pricing

Resources

About us

Contact us Login

Contact us Login

Clear

For a long time, getting the GPUs meant you were ready. Rack them, power them, cool them, and the dashboards go green. That was the whole test. It isn't anymore. Real AI workloads don’t ask whether the compute exists. It asks whether compute, networking, storage, scheduling, and recovery can hold together for two weeks straight, without stalling somewhere at the seams. AI factories keep every layer of the stack working in sync to turn power into tokens, reduce training run times from days to hours and accelerate token throughput. The bar keeps rising That list of supporting services only grows. Observability, health management, and the operations work in concert while jobs run in the background. And increasingly, some of the new jobs never really end. Due to the nature of agentic workloads, the end user doesn’t want to wait to see results . As agentic AI demand increases, so does token throughput. AI clouds cannot sacrifice reliability over speed. An agent calling tools or running a multi-step reasoning loop can spike demand at any hour, therefore teams must prepare their infrastructure to support variable demand. Delivering production AI Here's what that looks like in practice. A stalled job can constrain GPU resources and delay project timelines. A lost checkpoint means redoing work that already happened once, on the same valuable hardware. And underutilized resources are still running up the bill whether they're producing tokens or not. In this fast pace environment, teams cannot afford to lose productivity or underutilize compute resources. NVIDIA and CoreWeave partner to co-design a full-stack platform for leading model builders, enterprises, and AI natives to deliver production-ready AI faster. The work behind the partnership Teams that depend on these systems need the whole stack to behave as one before a workload goes live. Our teams work together from concept and design to stand up and validation of the latest systems to optimize performance and efficiency of the entire stack. NVIDIA is the accelerated computing platform, and CoreWeave provides the AI-native cloud software and operations layer that runs it in production. Together, we are making sure the full-stack behaves the same in a customer's hands as it did in validation. That work started in January 2026, when NVIDIA and CoreWeave expanded their collaboration to deploy Vera Rubin platform on CoreWeave using the NVIDIA DSX platform  DSX is NVIDIA's reference design for how AI factories are planned, built, and operated at maximum efficiency and profitability. The goal is deeper interoperability assuring more resilient and performance AI Factory capability. As a result of this collaboration, CoreWeave recently shared early performance benchmarks: Vera Rubin NVL72 generated 10x tokens-per-second per megawatt compared to NVIDIA GB200 NVL72 on the same DeepSeek R1 workload. This milestone demonstrates how closely our teams work together to bring-up next generation AI compute. Track record, not talking points ‍ Two results made that case in public. Last year, CoreWeave was the first AI cloud provider to deploy NVIDIA GB300 NVL72 .  And in MLPerf Training v6.0, CoreWeave completed the DeepSeek-V3 671B training benchmark in approximately two minutes on the largest NVIDIA GB300 NVL72 cluster in the round.  That track record isn't accidental. It's what happens when every layer of the stack is designed to work  together before a customer depends on it, and proven under real workloads before it ships. That's the shape of the full collaboration. Co-designed, validated before deployment, operated with confidence at scale, and carried into the next generation of hardware. Here's how we co-design and validate that stack, before turning to how it operates at scale and carries into the next generation in part two. ‍

1. Co-designed for production AI PFLOPs may appear to be the obvious indication of performance, however it doesn’t tell the full picture. Extreme co-design across GPUs, CPUs,  networking, and storage is required to achieve the highest performance possible. Every layer is designed in step with every other layer, before any of it reaches a customer. Planned before ground breaks The NVIDIA DSX Platform brings together land, power, and cloud partners, CoreWeave among them, to plan AI factories before they're built. That's extreme co-design: the whole ecosystem agreeing on how every layer fits together before a customer ever depends on it. Before we break ground on a new data center, networking, cooling, storage, and software are planned together to ensure end-to-end efficiency and performance. NVIDIA backs that up with hardware and software references designs built specifically to solve infrastructure problems at scale. Integration failures that only surface at large scale get caught by our engineers before onboarding, not in a customer's production run. Proven in Exemplar Cloud ‍ NVIDIA Exemplar Cloud is where performance at scale gets tested. It's a validation initiative that measures performance per TCO against real workloads, not synthetic ones CoreWeave was among the first to achieve Exemplar Cloud status for training on NVIDIA GB200 NVL72. The validation ran on a standard NVIDIA GB200 NVL72 cluster of 576 Blackwell GPUs, interconnected with NVIDIA Quantum-2 InfiniBand and performance-optimized by CoreWeave Mission Control, and it met NVIDIA's performance reference targets for large-scale training. Meaning the result comes from a trusted configuration customers will actually use rather than a one-off build. On the inference side, CoreWeave achieved Exemplar Cloud validation for inference on Blackwell GPUs, across DeepSeek-R1, Llama 3.3, and GPT-OSS, using NVIDIA TRT-LLM and SGLang backends with NVIDIA Dynamo for multi-node serving.  Same discipline, applied to the workload that runs every day, not just the one that runs once to set a benchmark. Carried into Next Generations ‍ That discipline doesn't reset with each new generation....

Excerpt shown — open the source for the full document.

Additional captured pages

© Copyright CoreWeave 2025. All rights reserved. CoreWeave, its logo, and coreweave.com are trademarks of CoreWeave, registered worldwide.This information is provided “as is” without any warranty, express or implied. This document is current as of the initial date of publication...

CV/ CoreWeave Supplier Code of Conduct Date of last review /update: November 2025 CoreWeave Supplier Spirit & Code of Conduct At CoreWeave, we have set the highest possible standards for the way we conduct business, and we expect that all of our Suppliers will lawfully conduct...

**WHITEPAPER** The infrastructure moment in AI Defining the Essential Cloud for AI © Copyright CoreWeave 2025. All rights reserved. CoreWeave, its logo, and coreweave.com are trademarks of CoreWeave,...

Notability

notability 5.0/10

Substantive industry blog post on AI validation