Benchmarks — who reports what
Which models self-report which benchmarks, parsed from their model cards. 0 benchmarks reported by 3+ models, across 0 models.
Not a leaderboard. These are each lab’s own reported numbers, measured under its own harness, config, prompt, and date. Versions and metrics differ — a higher number here does not mean a better model. Read them as “what each lab claims,” not a ranking.
No parsed benchmark data yet.