Aya Expanse Connecting Our World
Captured source
source ↗North Mini Code. Cohere's first model for developers.
Oct 24, 2024
7 minutes read
Aya Expanse: Connecting our world
Cohere For AI launches Aya Expanse, a state-of-the-art multilingual family of models to help close the language gap with AI.
Today, Cohere For AI, Cohere’s research arm, is proud to announce Aya Expanse, a family of highly performant multilingual models that excels across 23 languages and outperforms other leading open-weights models.
We are releasing Aya Expanse as both 8 and 32 billion open-weights models available on Kaggle and Hugging Face, as part of our continued commitment to multilingual research and to accelerate the frontier for multilingual AI. The 8 billion parameters model makes breakthroughs more accessible to researchers worldwide, and our 32 billion parameters model offers state-of-the-art multilingual capabilities.
Aya Expanse marks an important step to expand high-quality coverage of languages in LLMs. Since we first launched the Aya initiative two years ago, we have collaborated with over 3,000 researchers from 119 countries to expand cutting-edge multilingual research. This included releasing the Aya collection, the largest multilingual dataset collection to-date, with 513 million examples, and critical evaluation sets for multilingual performance and safety. It has also included the release of Aya-101, the most comprehensive multilingual model to-date covering 101 languages.

A Spotlight on Multilingual Research
At a larger scale, Aya Expanse 32B outperforms Gemma 2 27B, Mistral 8x22B, and Llama 3.1 70B, a model more than 2x its size, setting a new state-of-the-art for multilingual performance. We also released Aya Expanse 8B, which outperforms the leading open-weights models in its parameter class such as Gemma 2 9B, Llama 3.1 8B, and the recently released Ministral 8B with win rates ranging from 60.4% to 70.6%.
The improvements in Aya Expanse are the result of a sustained focus on expanding how AI serves languages around the world by rethinking the core building blocks of machine learning breakthroughs. Our research agenda for the last few years has included a dedicated focus on bridging the language gap, with several breakthroughs that were critical to the current recipe: data arbitrage, preference training for general performance and safety, and finally model merging. Below, we spotlight this research.
Advancing synthetic data for languages with limited data. The use of synthetic data – data generated by an expert or “teacher” model to train another model – has become increasingly central to the development of LLMs, particularly as model training has exhausted data sources. However, for multilingual data, especially with languages that are low-resource, there are few good examples of teacher models, creating an extra challenge to leveraging synthetic data.
We proposed a novel data sampling strategy that we term data arbitrage to avoid mode collapse, or the generation of “gibberish” when over relying on synthetic data. Data arbitrage takes inspiration from how humans learn by going to different teachers for different skills. If you want to learn to play piano, you’d go to a piano teacher. To learn to bake, you'd seek out a baking expert. We apply the same philosophy to AI, strategically selecting different “teacher” models based upon the data distribution to generate suitable synthetic data for multilingual capabilities.
Guiding models towards global preferences. Preference training is used in the late stages of model training, leveraging feedback from humans to guide the model toward what high-quality outputs look like. We think of it as the "final sparkle” in training an AI model. However, preference training and safety measures often overfit to harms prevalent in Western-centric datasets. Problematically, these safety protocols frequently fail to extend to multilingual settings. Our work is one of the first that extends preference training to a massively multilingual setting, accounting for different cultural and linguistic perspectives. We find this leads to large gains both in general performance and safety.
Model merging in a diverse world. The final step to our breakthrough in multilingual performance is our work on model merging – combining the weights of multiple candidate models at each stage to create more versatility and performance.

Putting together our research to build the best multilingual model in its class. We combined all of these techniques in one training recipe for Aya Expanse. Each of these techniques – from data arbitrage to merging and multilingual preference optimization – enable step-by-step improvement, leading to a significant gain...
Excerpt shown — open the source for the full document.
Notability
notability 8.0/10Multilingual model release with strong performance.