paper-with-me

홈 › Papers

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech

2026-03-17 · Omnilingual SONAR Team, João Maria Janeiro, Pere-Lluís Huguet Cabot, Ioannis Tsiamas, Yen Meng, Vivek Iyer, Guillem Ramírez, Loic Barrault, Belen Alastruey, Xiang "Tony" Cao, Yu-An Chung, Marta R. Costa-Jussa, David Dale, Kevin Heffernan, Jaehyeong Jo, Artyom Kozhevnikov, Alexandre Mourachko, Christophe Ropers, Holger Schwenk, Paul-Ambroise Duquenne arxiv

Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSONAR, a new family of omnilingual, cross-lingual and cross-modal sentence embedding models that natively embed text, speech, code, and mathematical expressions in a single semantic space, while delivering state-of-the-art downstream performance at the scale of thousands of languages, from high-resource to extremely low-resource varieties. To reach this scale without representation collapse, we use progressive training. We first learn a strong foundational space for 200 languages with an LLM-initialized encoder-decoder, combining token-level decoding with a novel split-softmax contrastive loss and synthetic hard negatives. Building on this foundation, we expand to several thousands language varieties via a two-stage teacher-student encoder distillation framework. Finally, we demonstrate the cross-modal extensibility of this space by seamlessly mapping 177 spoken languages into it. OmniSONAR halves cross-lingual similarity search error on the 200-language FLORES dataset and reduces error by a factor of 15 on the 1,560-language BIBLE benchmark. It also enables strong translation, outperforming NLLB-3B on multilingual benchmarks and exceeding prior models (including much larger LLMs) by 15 chrF++ points on 1,560 languages into English BIBLE translation. OmniSONAR also performs strongly on MTEB and XLCoST. For speech, OmniSONAR achieves a 43% lower similarity-search error and reaches 97% of SeamlessM4T speech-to-text quality, despite being zero-shot for translation (trained only on ASR data). Finally, by training an encoder-decoder LM, Spectrum, exclusively on English text processing OmniSONAR embedding sequences, we unlock high-performance transfer to thousands of languages and speech for complex downstream tasks.

📄 PDF Abstract BibTeX arXiv:2603.16606

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unified Vision-Language Modeling via Concept Space Alignment

2026-03-01 · Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk arxiv

We introduce V-SONAR, a vision-language embedding space extended from the text-only embedding space SONAR (Omnilingual Embeddings Team et al., 2026), which supports 1500 text languages and 177 speech languages. To constr…

Question AnsweringVideo CaptioningVideo Retrieval

Omnilingual MT: Machine Translation for 1,600 Languages

2026-03-17 · Omnilingual MT Team, Belen Alastruey, Niyati Bafna, Andrea Caciolai 외 arxiv

High-quality machine translation (MT) can scale to hundreds of languages, setting a high bar for multilingual systems. However, compared to the world's 7,000 languages, current systems still offer only limited coverage: …

Cross-Lingual TransferMachine Translation

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

2025-11-12 · Omnilingual ASR team, Gil Keren, Artyom Kozhevnikov, Yen Meng 외 arxiv

Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's 7,000+ languages remain unsupported, leaving thousands of long-tail languages behind. Expanding ASR coverage has been co…

Zero-shot GeneralizationSpeech Recognition

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

2026-06-25 · Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan, Astut Kurariya 외 arxiv

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, w…

Speech Recognition

A Sonar-Visual Dataset for Cross-Modal Underwater Robot Perception

2026-05-31 · Weitung Chen, Phil Tinn, Per Gunnar Auran, Martin Ludvigsen 외 arxiv

Underwater robots typically use both cameras and sonar for perception to leverage the rich semantic details of vision and the robust range measurements of acoustics. However, learning to map between these modalities via …