The Iol Ai Challenge An Open Challenge Towards Advancing Linguistic Reasoning 2026 08 18
Captured source
source ↗The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning Skip to content Products
Solutions
Resources
Blog
Research
Company
Sign in
Request a demo
Platform North
Enterprise-ready AI for business
Compass
Intelligent search and discovery
Models Command New
Generative language models
Transcribe New
Speech recognition model
North Mini Code New
Agentic coding model
Embed
Search and discovery model
Rerank
Semantic search ranking
Models Overview
Product Products Overview
Total Cost of AI Ownership
Pricing
Featured Command: High-performance generative AI models for real-world applications
Deploy Model Vault
Dedicated model inference platform
Private Deployments
On-prem or isolated VPCs
Security
Protect your data at every stage
See deployment options
By Industry Financial Services
Public Sector
Technology
Telecommunications
Energy and Utilities
Healthcare and Life Sciences
Manufacturing
Featured Model Vault provides fully-isolated, performant inference with Saas simplicity
Insights Customer Stories
For Developers Developers
Models Overview
Docs
Discord
LLM University
Connect Partners
Events
Webinars
Merch Store
Featured How CoreWeave used Cohere North to transform its customer support in 90 days
Blog
The latest news, launches, and insights
Read more
Cohere and the University of Toronto partner to advance responsible AI adoption at scale
Hardware-aware dynamic speculative decoding
Meet Cohere Transcribe Arabic
Research Cohere Labs
Cohere’s ML research lab
Explorations Future(s) of Work
How will AI change the way we work?
Aya Models
Multilingual AI at scale
All Papers
Initiatives Research Scholars
Finding the new generation of ML talent
Open Science Community
Championing global, open science
Catalyst Grants
Supporting impactful ML endeavors
Resources Blog
Hugging Face
Events
Featured The future of work debate has an evidence problem
About
Careers
Newsroom
Aug 18, 2026 The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning We hosted an open competition on unseen linguistic reasoning problems to measure and advance AI reasoning skills. The official IOL 2026 jury graded AI solutions, and one frontier model scored equivalent to a gold medal.
Read the paper
Authors
Eduardo Sánchez*, Rita Berrada*, Dan-Mircea Mirea*, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, and Julia Kreutzer
*First authors
Abstract
Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5\% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ${\sim}13$ points and understating strong ones.
Reasoning Evaluation Language Models
Related works
Research CIRCLE: A Framework for Evaluating AI from a Real-World Lens
Read
Research Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation
Read
Research SimMerge: Learning to Select Merge Operators from Similarity Signals
Read
Notability
notability 6.0/10Open challenge launch by Cohere, substantive but not a model release