WritingCohereCoherepublished Aug 18, 2026seen 1w

The Iol Ai Challenge An Open Challenge Towards Advancing Linguistic Reasoning 2026 08 18

Open original ↗

Captured source

source ↗

The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning Skip to content Products

Solutions

Resources

Blog

Research

Company

Sign in

Request a demo

Platform North

Enterprise-ready AI for business

Compass

Intelligent search and discovery

Models Command New

Generative language models

Transcribe New

Speech recognition model

North Mini Code New

Agentic coding model

Embed

Search and discovery model

Rerank

Semantic search ranking

Models Overview

Product Products Overview

Total Cost of AI Ownership

Pricing

Featured Command: High-performance generative AI models for real-world applications

Deploy Model Vault

Dedicated model inference platform

Private Deployments

On-prem or isolated VPCs

Security

Protect your data at every stage

See deployment options

By Industry Financial Services

Public Sector

Technology

Telecommunications

Energy and Utilities

Healthcare and Life Sciences

Manufacturing

Featured Model Vault provides fully-isolated, performant inference with Saas simplicity

Insights Customer Stories

For Developers Developers

Models Overview

Docs

Discord

LLM University

Connect Partners

Events

Webinars

Merch Store

Featured How CoreWeave used Cohere North to transform its customer support in 90 days

Blog

The latest news, launches, and insights

Read more

Cohere and the University of Toronto partner to advance responsible AI adoption at scale

Hardware-aware dynamic speculative decoding

Meet Cohere Transcribe Arabic

Research Cohere Labs

Cohere’s ML research lab

Explorations Future(s) of Work

How will AI change the way we work?

Aya Models

Multilingual AI at scale

All Papers

Initiatives Research Scholars

Finding the new generation of ML talent

Open Science Community

Championing global, open science

Catalyst Grants

Supporting impactful ML endeavors

Resources Blog

Hugging Face

Events

Featured The future of work debate has an evidence problem

About

Careers

Newsroom

Aug 18, 2026 The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning We hosted an open competition on unseen linguistic reasoning problems to measure and advance AI reasoning skills. The official IOL 2026 jury graded AI solutions, and one frontier model scored equivalent to a gold medal.

Read the paper

Authors

Eduardo Sánchez*, Rita Berrada*, Dan-Mircea Mirea*, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, and Julia Kreutzer

*First authors

Abstract

Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5\% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ${\sim}13$ points and understating strong ones.

Reasoning Evaluation Language Models

Related works

Research CIRCLE: A Framework for Evaluating AI from a Real-World Lens

Read

Research Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation

Read

Research SimMerge: Learning to Select Merge Operators from Similarity Signals

Read

Notability

notability 6.0/10

Open challenge launch by Cohere, substantive but not a model release