Frontier lab intelligence

Stored source text and lab context from public AI signals. Open a signal to read the captured source inside onlylabs; original links are secondary.

Get updates — new model intelligence in your inbox. No spam.
Signals / day · 30d365 this week · 1,000 total
Signal mix
Lead source

Building Multilingual Bridges 2026 09 10

Cohere multilingual blog post, substantive, no traction data.

Cohere · writing · cohere.comRead captured source

Deep reports

All reports
Three flagship cards, no leaderboardEval / RL-environmentsSell to labs

A cross-card read of GPT-5.6, GLM-5.2 and Kimi-K2-Thinking: the closed leader publishes zero capability comparisons while the two open labs publish incompatible ones — so the leaderboard everyone wants cannot be built from primary sources. That void is the eval product.

GPT-5.6 Preview System Card — Eval IntelligenceEval / RL-environments

Master benchmark inventory of the GPT-5.6 system card with web-verified cross-vendor comparison and the eval-shift Sankey — every claim tied to a page.

GLM-5.2 — the open agentic-frontier playEval / RL-environments

Zhipu’s MIT-licensed flagship is in the Opus/GPT-5.5/Gemini tier on agentic SWE while trailing on broad knowledge — and its own cross-vendor table shows why "the harness is the benchmark".

Kimi-K2-Thinking — the open agentic-research betEval / RL-environments

Moonshot’s open 1T-MoE thinking agent holds the top of agentic web research (BrowseComp SOTA, 200–300 tool calls without drift) while trailing closed frontier on raw coding — a distinct open frontier from GLM-5.2’s agentic engineering.

Evals at the frontier labs — what the hiring revealsGet hiredSell to labs

What 128 open eval-relevant roles reveal about frontier eval investment, for two audiences: people who want to get hired, and people who want to sell to the labs.

Infra & systems at the frontier — what the hiring revealsGet hiredSell to labs

What 715 open infrastructure/systems roles reveal about the GPU buildout — physical first — for two audiences: people who want to get hired in infra, and people who want to sell infra to the labs and neoclouds.

Safety & alignment at the frontier — what the hiring revealsGet hiredSell to labs

What 127 open safety/alignment/red-team roles reveal about the safety org (OpenAI out-hires Anthropic in raw count), for two audiences: get hired into safety, and sell safety tooling to the labs.

Human data & annotation at the frontier — what the hiring revealsGet hiredSell to labs

What 86 open human-data / annotation / data-quality roles reveal about the fuel layer — the clearest "sell to the labs" buy signal, since labs structurally buy data rather than build it.

In depth

All analysis
Amazon (Nova)Amazon (Nova)9h

Amazon's public signal this cycle is bifurcated. Press reporting indicates most flagship Nova models are being wound down and resources consolidated into Frontier Model Research under Pieter Abbeel, with a new flagship expected at re:Invent , while Amazon Science continues shippi

AnthropicAnthropic1d

Anthropic's September 2026 signals show a lab running three plays at once: pushing frontier capability into math and science ; hardening and publicly disclosing its safety/alignment machinery ; and industrializing compute, finance, and go-to-market ahead of a reported IPO . The w

DeepSeekDeepSeek1w

DeepSeek's center of gravity is shifting from flagship frontier models toward agentic productization and efficiency infrastructure. The dominant recent artifact is DeepSeek Harness ('dsh'), an MIT-licensed TypeScript agent harness built on an 'everything is a plugin' Cordis archi

CohereCohere1w

Cohere is running a security-first, enterprise-and-sovereign play rather than a consumer-scale model race. The dominant signals in this pack are commercialization and deployment buildout: a dense wave of forward-deployed engineering (FDE) hiring across Infrastructure, Agentic Pla

Baidu (ERNIE)Baidu (ERNIE)2w

Baidu is running two visible plays in this pack. First, a heavy open-source, deployment-and-interoperability push across the PaddlePaddle ecosystem — PaddlePaddle v3.0, PaddleOCR v3.x, PaddleX v3.x, and Paddle2ONNX v2.x — aimed at multi-hardware inference, serving, and model expo

234 loaded · 26,741 total