Rapidly scaling online storage to serve over 1 billion ChatGPT users
Captured source
source ↗Rapidly scaling online storage to serve over 1 billion ChatGPT users | OpenAI
September 11, 2026
Rapidly scaling online storage to serve over 1 billion ChatGPT users
How we adapted our application storage platform, Habitat, in Python to manage unprecedented growth. By Jon Lee, Chaomin Yu, and Ben Ries, Members of Technical Staff
Every OpenAI product depends on fast, reliable access to data, whether someone is logging in, checking their Codex settings, or starting a new conversation in ChatGPT. Each of those actions may require many separate data lookups before the product can respond. If those requests are slow, the product feels slow. If those requests fail, the product stops working entirely.
Habitat is the online storage platform we built so OpenAI products can quickly and reliably access needed information. Habitat now handles more than 70 million requests every second, supporting products used by over 1 billion people each week, across almost 40 geographic regions. Two years ago, Habitat started as a simple Python client-side library connected to a single database. Today, it’s a complex distributed system that serves more than 500 petabytes of data.
Figure 01 · What is Habitat? Online storage platform Habitat is the online storage platform we built so OpenAI products can quickly and reliably access needed information. Play
- Request
- Response
- Changes (CDC)
Clients
Online storage platform
Storage resources
- ChatGPT
- API
- Codex
- Internal services
- And more
Habitat
- Caching Caches
- ACL policies Authorization
- Placement & data residency Data residency
- Encryption Data security
- Isolation Multi-tenancy
- Rate limiting Request shaping
- Routing Schema lookup · Data residency
- Azure Cosmos DB Online storage
- Nanobase Online storage
- Valkey Caches
- Blob storage Storage resources
- CDC Services Change Data Capture
- Databricks
- Rockset
- Kafka
- And more
Building and operating infrastructure at this scale is no easy feat, but also not particularly challenging. What made our situation unique is the unprecedented rate at which we’ve had to scale to support staggering user growth and product demand while simultaneously building out a mature platform. Often, system engineers build for 10x scale, and hope for it to hold for a few years while preparing for the next 10x. In our case, we've grown more than 10x year-over-year for the last three years. As a result, building and operating Habitat has been a series of tactical decisions and sequencing: understanding each component at the lowest level to squeeze as much juice out of our existing stack, while fending off storage and compute capacity crunches to buy time for foundational investments.
- 70M+
requests per second
- 1B+
people each week
- 500 PB+
data
As OpenAI grew, Habitat had to grow with it: first by becoming reliable enough for mission-critical product traffic, then fast enough for global users, and finally, to deftly operate at massive scale. This post is the first in a two-part series on how we scaled online storage. In this post, we’ll share how Habitat evolved, why we turned it from a library into a service, and how we stretched a service written in an uncommon serving stack language—Python—into a reliable storage platform layer.
In a future post, we’ll go into detail about how we made multi-tenancy reliability at scale, our layered strategy for optimizing read performance, and how we scaled our partnership with Azure Cosmos DB to reliably handle unprecedented demand.
What is Habitat?
Habitat started from a simple idea: product engineers shouldn’t need to think about database management. Habitat began in mid-2024 as a small Python library that interacted with ChatGPT’s main server. It supported a small set of operations that mapped under the hood to the database application, Azure Cosmos DB.
The library’s job was to give product teams a simple way to store and retrieve data without needing to master the underlying details. Habitat took care of the necessary work: figuring out what kind of data was involved, where it should come from (or go), whether the request was allowed, and so on.
Product engineers need not concern themselves with schema lookup, routing, authorization, encryption, serialization, request shaping, and connection pooling. They didn’t even need to consider where the data comes from: Azure Cosmos DB, caches, or other types of storage.
Figure 02 · Habitat service Simplified Habitat request flow By decoupling the storage logic into a standalone service, we established a single point of control for deployments, observability, and platform enhancements. Play
- Request
- Response
Client
OpenAI
Azure Cosmos DB
- Habitat client sdk
- envoy
- habitat-service process 1
- habitat-service process 2
- habitat-service process 3
- habitat-envoy
- habitat-cosmos-db-us0
- habitat-cosmos-db-us1
- habitat-cosmos-db-eu0
This Python library worked well and Habitat saw rapid adoption among product engineers at OpenAI, despite no concerted central push away from using self-serve Postgres and Azure Cosmos DB.
As product needs evolved, it was even easy for product developers to add to the shared library support for features like client-side caching, compression, or encryption.
Build a service to better support multiple, complex products
By the middle of 2025, Habitat had reached its limits as a client-side implementation. As the Habitat layer had grown more complex and OpenAI’s services count increased, backward-compatible protocol changes had become infeasible.
In one instance, we wanted to reduce the blast radius of any single region outage for our most critical data sets by migrating them to a set of regionally distributed Azure Cosmos DB accounts. Making this change required introducing extra routing logic into the client, disabled behind a feature flag, ensuring it rolled out to all clients, and then enabling the feature flag.
Coordinating deployments across dozens of services and working with each team to roll it out took days. Before enabling this, we realized we wanted to introduce some shadowing to ensure the sharding logic would be correct. That took another couple of days to roll out. A bug fix for something we realized was incorrect? Another couple of days. Eventually, we were ready to enable the flag, only for one of the teams to roll back their service for unrelated reasons to a previously buggy client, causing the outage...
Excerpt shown — open the source for the full document.
Notability
notability 5.0/10Substantive engineering post, low traction
OpenAI has a writing signal matching data demand, infrastructure.