Three flagship cards, no leaderboard
Charts
Three flagship cards, no leaderboard
A cross-card read of the three flagship releases onlylabs has deep-reported — GPT-5.6 (OpenAI, closed), GLM-5.2 (Zhipu, open) and Kimi-K2-Thinking (Moonshot, open). The question an eval founder actually has — "which is best?" — these three documents are structurally unable to answer. That inability is the finding.
0. The setup
Three frontier-tier models shipped their model cards. A buyer, a candidate, or a journalist wants one table that ranks them. You cannot build it from these cards — not because the labs hid the data, but because the three documents are three different genres measuring three different things three different ways. This report is the cross-card synthesis: what each one discloses, why the columns don't merge, and what the shape of the disclosure tells you about each lab.
1. Three disclosure regimes
| Flagship card | License | Cross-vendor capability table? | What the card foregrounds |
|---|---|---|---|
| GPT-5.6 (Sol / Terra / Luna) system card | Closed | None — zero competitor capability columns; GPQA, HLE, SWE-bench appear only as substrates inside other evals | Preparedness thresholds (bio / chem / cyber), refusal & jailbreak rates, CoT monitorability |
| GLM-5.2 | MIT (open) | Yes — vs Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek-V4-Pro, MiniMax, Qwen | Agentic engineering — FrontierSWE, Terminal-Bench, SWE-bench Pro, SWE-Marathon |
| Kimi-K2-Thinking | Modified-MIT (open) | Yes — vs GPT-5 (High), Claude Sonnet 4.5, DeepSeek-V3.2, Grok-4 | Agentic research — BrowseComp, Seal-0, Frames + 200–300-step tool horizon |
The single most underrated fact about the current frontier: the closed leader publishes no capability leaderboard at all. OpenAI's GPT-5.6 card is ~70 pages of safety and preparedness eval machinery. The standard capability benchmarks a ranking would need — GPQA, HLE, SWE-bench Verified — are present only as substrates: the raw material inside CoT-Control (>13,000 tasks built from them) or PostTrainBench (objectives an agent trains toward), never as a reported GPT-5.6 score. Meanwhile the two open Chinese labs publish full cross-vendor tables. The disclosure gradient runs opposite to the openness gradient.
2. Why the columns don't merge — one benchmark, eight numbers
Take HLE (Humanity's Last Exam), the one benchmark that nominally appears across all three ecosystems. Line up every HLE figure these cards (and the web-verified comparands in the GPT-5.6 report) actually report:
- GPT-5.6: no number — HLE is a substrate, not a score.
- GLM-5.2: 40.5 (text-only) · 54.7 (with tools)
- Kimi-K2-Thinking: 23.9 (no tools) · 44.9 (with tools) · 51.0 (heavy)
- Gemini 3.1 Pro (web-verified context): 44.4 (no tools) · 51.4 (with tools)
- DeepSeek-V3.2 (web-verified context): 25.1 (text-only, no tools)
That is eight numbers under one label — and they range from 23.9 to 54.7 on the same exam, driven almost entirely by tool-setting and subset (text-only vs full, no-tools vs with-tools vs "heavy"). A single "HLE" column in a leaderboard would silently average a no-tools score against a heavy-agentic one. The chart below makes the spread visible — and it is explicitly not a ranking.
The same trap repeats on coding. "SWE-bench" is at least three different sets here: GPT-5.6 reports SWE-bench Verified only as a substrate; GLM-5.2 reports SWE-bench Pro (62.1, via OpenHands); Kimi reports SWE-bench Verified (71.3, with tools). And BrowseComp — Kimi's headline win (60.2) — appears on only one of the three cards, so the three literally cannot be ranked on agentic web research from their own disclosures.
3. What each card won't let you conclude
- You cannot rank GPT-5.6 against the open pair on capability. It declines to publish the numbers. Any "GPT-5.6 vs GLM-5.2" capability table on the internet is stitched from third-party runs, not the card.
- You cannot merge the two open tables. GLM-5.2 vs Opus-4.8/GPT-5.5/Gemini-3.1; Kimi vs GPT-5/Sonnet-4.5/Grok-4 — different competitor sets, different reference generations, different harnesses (OpenHands vs the Kimi agent), different judges. GLM's own footnotes admit each row uses a different scaffold/context-window/judge.
- You cannot trust a bare benchmark name. As §2 shows, "HLE" or "SWE-bench" without the setting (tools / heavy / subset / scaffold) is not a rankable number. The harness is the benchmark.
4. What the shape of disclosure tells you (two audiences)
The disclosure regime is itself a signal — and it's the most actionable thing here for the two audiences onlylabs serves.
If you sell to the labs:
- OpenAI's card is a shopping list for safety/eval vendors. ~70 pages of preparedness evals — bio/chem thresholds, CVE-Bench, automated red-teaming at >700k A100e GPU-hrs, CoT monitors, Apollo sandbagging audits, METR time-horizon. That machinery is built and bought. If your product is red-teaming, monitoring, preparedness measurement, or third-party capability eval, the GPT-5.6 card is the buyer's spec sheet.
- The open Chinese labs are racing on a public capability table — and several rows are measured by independent third parties (FrontierSWE by Proximal, SWE-Marathon by Abundant AI). That third-party-eval pattern is the emerging business: the labs increasingly outsource the scoring they want to be credible.
If you want to get hired:
- The disclosure shape matches the hiring shape we see in the persona reports: OpenAI's safety/preparedness eval apparatus (and the headcount behind it — OpenAI out-hires Anthropic on raw safety count) is exactly what a 70-page safety card requires. The open labs' capability-table racing maps to agentic-coding and agentic-search eval roles.
The product gap, stated plainly: there is no vendor-neutral, same-harness, cross-lab capability measurement these three cards could be poured into. The closed leader won't disclose; the open labs disclose incompatibly. That void — a leaderboard you cannot build from primary sources — is the eval product. Not another benchmark. The harness that runs the existing benchmarks the same way across every model, with the setting (tools / subset / scaffold / judge / context window) as a first-class, reported dimension.
Every number above is the labs' own self-reported figure or a web-verified public comparand, carried over from the three linked deep reports. None of it is apples-to-apples — that is the entire point. See the /benchmarks "who reports what" view for the underlying self-reported scores.