paper-with-me

Papers

Aya 23: Open Weight Releases to Further Multilingual Progress

2024-05-23 · Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Aidan Gomez, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, Sara Hooker

This technical report introduces Aya 23, a family of multilingual language models. Aya 23 builds on the recent release of the Aya model (\"Ust\"un et al., 2024), focusing on pairing a highly performant pre-trained model with the recently released Aya collection (Singh et al., 2024). The result is a powerful multilingual large language model serving 23 languages, expanding state-of-art language modeling capabilities to approximately half of the world's population. The Aya model covered 101 languages whereas Aya 23 is an experiment in depth vs breadth, exploring the impact of allocating more capacity to fewer languages that are included during pre-training. Aya 23 outperforms both previous massively multilingual models like Aya 101 for the languages it covers, as well as widely used models like Gemma, Mistral and Mixtral on an extensive range of discriminative and generative tasks. We release the open weights for both the 8B and 35B models as part of our continued commitment for expanding access to multilingual progress.

📄 PDF Abstract BibTeX arXiv:2405.15032

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Massively Multilingual Word Embeddings

2016-02-05 · Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample 외

We introduce new methods for estimating and evaluating embeddings of words in more than fifty languages in a single shared embedding space. Our estimation methods, multiCluster and multiCCA, use dictionaries and monoling…

Multilingual Word EmbeddingsText CategorizationWord Embeddings

EuroLLM: Multilingual Language Models for Europe

2024-09-24 · Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro 외

The quality of open-weight LLMs has seen significant improvement, yet they remain predominantly focused on English. In this paper, we introduce the EuroLLM project, aimed at developing a suite of open-weight multilingual…

Machine Translation

SwissAdmin: A multilingual tagged parallel corpus of press releases

2014-05-01 · LREC 2014 5 · Yves Scherrer, Luka Nerima, Lorenza Russo, Maria Ivanova 외

SwissAdmin is a new multilingual corpus of press releases from the Swiss Federal Administration, available in German, French, Italian and English. We provide SwissAdmin in three versions: (i) plain texts of approximately…

Language IdentificationSentence

PRiSM: Benchmarking Phone Realization in Speech Models

2026-01-20 · Shikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim, Kwanghee Choi 외 arxiv

Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only…

Extracting the Structure of Press Releases for Predicting Earnings Announcement Returns

2025-09-29 · Yuntao Wu, Ege Mert Akin, Charles Martineau, Vincent Grégoire 외 arxiv

We examine how textual features in earnings press releases predict stock returns on earnings announcement days. Using over 138,000 press releases from 2005 to 2023, we compare traditional bag-of-words and BERT-based embe…