WritingBasetenBasetenpublished Jul 30, 2026seen Jul 31

Inkling Small Available On Baseten

Open original ↗

Captured source

source ↗
published Jul 30, 2026seen Jul 31captured Jul 31http 200method plain

Inkling Small available on Baseten Kimi K3 is here. Try it now

changelog / post

Inkling Small available on Baseten

Jul 30, 2026 Go back

Inkling Small is now available through Baseten Model APIs. Send requests to thinkingmachines/inkling-small through our OpenAI-compatible endpoint with your Baseten API key. Dedicated deployments are also available for larger workloads. Inkling Small is Thinking Machines Lab’s open-weights, 276B-parameter mixture-of-experts model with 12B active parameters. It retains Inkling’s 1M-token context window, native text, image, and audio inputs, tool calling, structured outputs, and controllable reasoning in a smaller model designed for workloads where latency and inference cost matter. Thinking Machines reports comparable performance to Inkling at roughly one-quarter its size. 1 curl https: //inference.baseten.co/v1/chat/completions \ 2 -H "Authorization: Bearer $BASETEN_API_KEY" \ 3 -H "Content-Type: application/json" \ 4 -d '{ 5 "model" : "thinkingmachines/inkling-small" , 6 "messages" : [ 7 { 8 "role" : "user" , 9 "content" : "Compare sparse and dense transformer architectures." 10 } 11 ], 12 "reasoning_effort" : "medium" 13 }' For supported reasoning settings and multimodal request examples, see the docs .

Explore Baseten today Start deploying Talk to an engineer

Popular models Kimi K3

GLM-5.2 Fast

Inkling

GLM-5.2

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Popular models Kimi K3

GLM-5.2 Fast

Inkling

GLM-5.2

Whisper Large V3

NVIDIA Nemotron 3 Ultra

Explore all

Notability

notability 3.0/10

Routine model availability announcement.