Kimi-K2-Thinking — the open agentic-research bet

Primary source ↗ · Synthesized by the onlylabs Content Studio agent (Claude Code) · web-verified

Charts

BrowseComp (agentic web research) — K2 Thinking leads (self-reported)
Kimi K2 Thinking (open)60.2GPT-5 (High)54.9DeepSeek-V3.240.1Claude Sonnet 4.524.1K2 0905 (prior)7.4
SWE-bench Verified — where K2 trails the closed leaders (self-reported)
Claude Sonnet 4.577.2GPT-5 (High)74.9Kimi K2 Thinking (open)71.3K2 0905 (prior)69.2DeepSeek-V3.267.8

Kimi-K2-Thinking — the open agentic-research bet

An eval-founder read of moonshotai/Kimi-K2-Thinking (Moonshot AI, modified-MIT, released 2025-11-04; tech blog). Numbers are the card's own self-reported figures — the §1 table is the cross-vendor comparison set; a few heavy / w/ python variants cited in the text are additional settings the same card reports. onlylabs signal.


0. What this is


1. The benchmark profile (self-reported, by setting)

The setting column matters as much as the score — tool use swings results enormously.

BenchmarkSettingK2 ThinkingGPT-5 (High)Sonnet 4.5 (Thinking)DeepSeek-V3.2Grok-4K2 0905
Agentic search
BrowseCompw/ tools60.254.924.140.17.4
Seal-0w/ tools56.351.4*53.4*38.5*25.2
Framesw/ tools87.086.0*85.0*80.2*58.1
Reasoning
HLE (text)w/ tools44.941.7*32.0*20.3*41.021.7
HLE (text)no tools23.926.319.8*19.825.47.9
AIME25no tools94.594.687.089.391.751.0
GPQAno tools84.585.783.479.987.574.2
Coding
SWE-bench Verifiedw/ tools71.374.977.267.869.2
LiveCodeBench V6no tools83.187.0*64.0*74.156.1*
Terminal-Benchsim. tools47.143.851.037.744.5
Knowledge
MMLU-Prono tools84.687.187.585.081.9
HealthBenchno tools58.067.244.246.943.8

2. Where it leads, where it trails


3. Comparability gotchas (the setting is the benchmark)


4. What it means for an eval / RL-environments founder

1. *Open agentic research is a distinct frontier from agentic engineering. Where GLM-5.2 bets on long-horizon coding, K2 Thinking bets on long-horizon browsing/search (BrowseComp, Seal-0, Frames) — and an open model now holds the top of that category. Two different open-weight agentic frontiers, two different eval stacks. 2. The benchmark that matters is a horizon benchmark. "200–300 sequential tool calls without drift" is the claim the field can't yet measure with a single number — the gap between a 50-step and a 300-step agent is exactly where current evals are thinnest. That's the eval product to build. 3. Efficiency is now a first-class eval axis.* Native-INT4-at-no-quality-loss reframes "score per token" into "score per dollar/latency." A vendor-neutral harness that reports cost/latency alongside accuracy would capture what these cards increasingly compete on.