Frontier lab intelligence

Stored source text and lab context from public AI signals. Open a signal to read the captured source inside onlylabs; original links are secondary.

Get updates — new model intelligence in your inbox. No spam.
Signals / day · 30d336 this week · 1,000 total
Signal mix
Lead source

Claude Opus 5

Anthropic · writing · anthropic.com · 845 HNRead captured source

Deep reports

All reports
Three flagship cards, no leaderboardEval / RL-environmentsSell to labs

A cross-card read of GPT-5.6, GLM-5.2 and Kimi-K2-Thinking: the closed leader publishes zero capability comparisons while the two open labs publish incompatible ones — so the leaderboard everyone wants cannot be built from primary sources. That void is the eval product.

GPT-5.6 Preview System Card — Eval IntelligenceEval / RL-environments

Master benchmark inventory of the GPT-5.6 system card with web-verified cross-vendor comparison and the eval-shift Sankey — every claim tied to a page.

GLM-5.2 — the open agentic-frontier playEval / RL-environments

Zhipu’s MIT-licensed flagship is in the Opus/GPT-5.5/Gemini tier on agentic SWE while trailing on broad knowledge — and its own cross-vendor table shows why "the harness is the benchmark".

Kimi-K2-Thinking — the open agentic-research betEval / RL-environments

Moonshot’s open 1T-MoE thinking agent holds the top of agentic web research (BrowseComp SOTA, 200–300 tool calls without drift) while trailing closed frontier on raw coding — a distinct open frontier from GLM-5.2’s agentic engineering.

Evals at the frontier labs — what the hiring revealsGet hiredSell to labs

What 128 open eval-relevant roles reveal about frontier eval investment, for two audiences: people who want to get hired, and people who want to sell to the labs.

Infra & systems at the frontier — what the hiring revealsGet hiredSell to labs

What 715 open infrastructure/systems roles reveal about the GPU buildout — physical first — for two audiences: people who want to get hired in infra, and people who want to sell infra to the labs and neoclouds.

Safety & alignment at the frontier — what the hiring revealsGet hiredSell to labs

What 127 open safety/alignment/red-team roles reveal about the safety org (OpenAI out-hires Anthropic in raw count), for two audiences: get hired into safety, and sell safety tooling to the labs.

Human data & annotation at the frontier — what the hiring revealsGet hiredSell to labs

What 86 open human-data / annotation / data-quality roles reveal about the fuel layer — the clearest "sell to the labs" buy signal, since labs structurally buy data rather than build it.

In depth

All analysis
DeepSeekDeepSeek3w

DeepSeek is executing a two-pronged strategy in mid-2026: aggressively optimizing inference economics through open-source speculative decoding infrastructure (DSpark, DeepSpec, Eagle3) while simultaneously expanding into the agentic coding product market with a new Code Harness t

CohereCohere3w

Cohere is transitioning from a research-forward multilingual lab into a security-first enterprise platform company. The evidence shows a lab simultaneously pushing research on MoE architectures and synthetic data while executing an aggressive commercialization buildout centered o

ByteDance (Doubao/Seed)ByteDance (Doubao/Seed)3w

ByteDance Seed is executing a deliberate multi-frontier strategy: shipping production models through Volcano Engine (Doubao 2.1 Pro, Seedance 2.5) while simultaneously open-sourcing a broad portfolio of research artifacts across LLM reasoning, multimodal understanding, video, bio

AnthropicAnthropic3w

Anthropic is transitioning from a model-research lab into a vertically integrated AI product company with global commercial ambitions. The evidence shows simultaneous acceleration across four fronts: (1) new model launches (Sonnet 5, Fable 5, Mythos 5) alongside the lifting of US

Amazon (Nova)Amazon (Nova)3w

Amazon is building a vertically integrated AI stack that runs from custom silicon and formally verified cloud infrastructure through foundation models, agent frameworks, and domain-specific evaluation tooling. The evidence reveals a three-pillar strategy: (1) infrastructure diffe

275 loaded · 22,324 total