Building Multilingual Bridges 2026 09 10
Captured source
source ↗Building Multilingual Bridges Skip to content AI for Empowerment: Your freedom. Your focus. See how AI gives you more time for what truly moves you. Explore now
Products
Solutions
Resources
Blog
Research
Company
Sign in
Request a demo
Platform North
Enterprise-ready AI for business
Compass
Intelligent search and discovery
Models Command
Generative language models
Transcribe New
Speech recognition model
North Small Translate NEW
Machine translation model
Parse New
Document parsing model
Embed
Semantic representation model
Rerank
Retrieval optimization model
Models Overview
Product Products Overview
Total Cost of AI Ownership
Pricing
Featured Command: High-performance generative AI models for real-world applications
Deploy Model Vault
Dedicated model inference platform
Private Deployments
On-prem or isolated VPCs
Security
Protect your data at every stage
See deployment options
By Industry Financial Services
Public Sector
Technology
Telecommunications
Energy and Utilities
Healthcare and Life Sciences
Manufacturing
Featured Model Vault provides fully-isolated, performant inference with Saas simplicity
Insights Customer Stories
For Developers Developers
Models Overview
Docs
Discord
LLM University
Connect Partners
Events
Webinars
Merch Store
Featured How CoreWeave used Cohere North to transform its customer support in 90 days
Blog
The latest news, launches, and insights
Read more
The state of sovereign AI adoption in 2026
Cohere and the University of Waterloo launch partnership to strengthen Canada’s AI talent pipeline
Introducing North Automations: Intelligent workflow orchestration
Research Cohere Labs
Cohere’s ML research lab
Explorations Future(s) of Work
How will AI change the way we work?
Aya Models
Multilingual AI at scale
All Papers
Initiatives Research Scholars
Finding the new generation of ML talent
Open Science Community
Championing global, open science
Catalyst Grants
Supporting impactful ML endeavors
Resources Blog
Hugging Face
Events
Featured The future of work debate has an evidence problem
About
Careers
Newsroom
Sep 10, 2026 Building Multilingual Bridges Data Mixing as the Pillar of Generalization for In-Language Reasoning
Optimized data mixtures including multilingual reasoning, non-reasoning and English reasoning data yield superior in-language reasoning. We build Tiny Aya L2-Thinker, a RLM supporting in-language reasoning on over 60 languages without sacrificing performance, the broadest coverage to-date for multilingual RLMs.
Read the paper
Authors
Mehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza, Alexandre Bérard, Thomas Euyang, Marzieh Fadaee, Julia Kreutzer
Abstract
Reasoning language models have substantially advanced on a variety of complex tasks, yet the capability remains overwhelmingly English-centric: even when prompted in another language, models predominantly reason in English. This is inaccessible for non-English-speaking users, risks losing the intent of the original question, and forgoes knowledge more readily expressed in the target language. In this work, we advance L2 reasoning , the ability of a model to reason consistently in the language of the user's prompt, thus building an in-language bridge between the prompt and the answer. We approach this problem from a data-centric angle, investigating how to optimize data composition and scheduling in SFT for reasoning generalization. Building on top of a massively multilingual 3.35B base model, we achieve above 95% L2 reasoning rate across 60 languages on 5 benchmarks spanning math, commonsense reasoning, instruction following and open-ended generation, while keeping performance strong. We show the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone. These findings indicate that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing and without requiring reasoning supervision in every target language. We release our model weights and multilingual reasoning data to support further research on accessible, in-language reasoning.
Reasoning multilingual
Related works
Research CALIBER: Calibrating confidence before and after reasoning in language models
Read
Research The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
Read
Research Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation
Read
Notability
notability 6.0/10Cohere multilingual blog post, substantive, no traction data.