Three flagship cards, no leaderboard

Primary source ↗ · Synthesized by the onlylabs Content Studio agent (Claude Code) · web-verified

Charts

One "HLE" label, eight numbers (23.9–54.7) — NOT a ranking; the point is incompatibility of setting
GLM-5.2 · text40.5GLM-5.2 · tools54.7Kimi · none23.9Kimi · tools44.9Kimi · heavy51Gemini 3.1 · none44.4Gemini 3.1 · tools51.4DeepSeek 3.2 · text25.1
Distinct external vendors each flagship card benchmarks itself against (self-reported) — the closed leader: zero
GLM-5.2 (open)6Kimi K2 Thinking (open)4GPT-5.6 (closed)0

Three flagship cards, no leaderboard

A cross-card read of the three flagship releases onlylabs has deep-reported — GPT-5.6 (OpenAI, closed), GLM-5.2 (Zhipu, open) and Kimi-K2-Thinking (Moonshot, open). The question an eval founder actually has — "which is best?" — these three documents are structurally unable to answer. That inability is the finding.


0. The setup

Three frontier-tier models shipped their model cards. A buyer, a candidate, or a journalist wants one table that ranks them. You cannot build it from these cards — not because the labs hid the data, but because the three documents are three different genres measuring three different things three different ways. This report is the cross-card synthesis: what each one discloses, why the columns don't merge, and what the shape of the disclosure tells you about each lab.


1. Three disclosure regimes

Flagship cardLicenseCross-vendor capability table?What the card foregrounds
GPT-5.6 (Sol / Terra / Luna) system cardClosedNone — zero competitor capability columns; GPQA, HLE, SWE-bench appear only as substrates inside other evalsPreparedness thresholds (bio / chem / cyber), refusal & jailbreak rates, CoT monitorability
GLM-5.2MIT (open)Yes — vs Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek-V4-Pro, MiniMax, QwenAgentic engineering — FrontierSWE, Terminal-Bench, SWE-bench Pro, SWE-Marathon
Kimi-K2-ThinkingModified-MIT (open)Yes — vs GPT-5 (High), Claude Sonnet 4.5, DeepSeek-V3.2, Grok-4Agentic research — BrowseComp, Seal-0, Frames + 200–300-step tool horizon

The single most underrated fact about the current frontier: the closed leader publishes no capability leaderboard at all. OpenAI's GPT-5.6 card is ~70 pages of safety and preparedness eval machinery. The standard capability benchmarks a ranking would need — GPQA, HLE, SWE-bench Verified — are present only as substrates: the raw material inside CoT-Control (>13,000 tasks built from them) or PostTrainBench (objectives an agent trains toward), never as a reported GPT-5.6 score. Meanwhile the two open Chinese labs publish full cross-vendor tables. The disclosure gradient runs opposite to the openness gradient.


2. Why the columns don't merge — one benchmark, eight numbers

Take HLE (Humanity's Last Exam), the one benchmark that nominally appears across all three ecosystems. Line up every HLE figure these cards (and the web-verified comparands in the GPT-5.6 report) actually report:

That is eight numbers under one label — and they range from 23.9 to 54.7 on the same exam, driven almost entirely by tool-setting and subset (text-only vs full, no-tools vs with-tools vs "heavy"). A single "HLE" column in a leaderboard would silently average a no-tools score against a heavy-agentic one. The chart below makes the spread visible — and it is explicitly not a ranking.

The same trap repeats on coding. "SWE-bench" is at least three different sets here: GPT-5.6 reports SWE-bench Verified only as a substrate; GLM-5.2 reports SWE-bench Pro (62.1, via OpenHands); Kimi reports SWE-bench Verified (71.3, with tools). And BrowseComp — Kimi's headline win (60.2) — appears on only one of the three cards, so the three literally cannot be ranked on agentic web research from their own disclosures.


3. What each card won't let you conclude


4. What the shape of disclosure tells you (two audiences)

The disclosure regime is itself a signal — and it's the most actionable thing here for the two audiences onlylabs serves.

If you sell to the labs:

If you want to get hired:

The product gap, stated plainly: there is no vendor-neutral, same-harness, cross-lab capability measurement these three cards could be poured into. The closed leader won't disclose; the open labs disclose incompatibly. That void — a leaderboard you cannot build from primary sources — is the eval product. Not another benchmark. The harness that runs the existing benchmarks the same way across every model, with the setting (tools / subset / scaffold / judge / context window) as a first-class, reported dimension.


Every number above is the labs' own self-reported figure or a web-verified public comparand, carried over from the three linked deep reports. None of it is apples-to-apples — that is the entire point. See the /benchmarks "who reports what" view for the underlying self-reported scores.