paper-with-me

홈 › Papers

Towards Robust Speech Representation Learning for Thousands of Languages

2024-06-30 · William Chen, Wangyou Zhang, Yifan Peng, Xinjian Li, Jinchuan Tian, Jiatong Shi, Xuankai Chang, Soumi Maiti, Karen Livescu, Shinji Watanabe

Self-supervised learning (SSL) has helped extend speech technologies to more languages by reducing the need for labeled data. However, models are still far from supporting the world's 7000+ languages. We propose XEUS, a Cross-lingual Encoder for Universal Speech, trained on over 1 million hours of data across 4057 languages, extending the language coverage of SSL models 4-fold. We combine 1 million hours of speech from existing publicly accessible corpora with a newly created corpus of 7400+ hours from 4057 languages, which will be publicly released. To handle the diverse conditions of multilingual speech data, we augment the typical SSL masked prediction approach with a novel dereverberation objective, increasing robustness. We evaluate XEUS on several benchmarks, and show that it consistently outperforms or achieves comparable results to state-of-the-art (SOTA) SSL models across a variety of tasks. XEUS sets a new SOTA on the ML-SUPERB benchmark: it outperforms MMS 1B and w2v-BERT 2.0 v2 by 0.8% and 4.4% respectively, despite having less parameters or pre-training data. Checkpoints, code, and data are found in https://www.wavlab.org/activities/2024/xeus/.

📄 PDF Abstract BibTeX arXiv:2407.00837

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSelf-Supervised LearningSpeech Representation Learning

Similar Papers 제목 키워드 기반

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech

2026-03-17 · Omnilingual SONAR Team, João Maria Janeiro, Pere-Lluís Huguet Cabot, Ioannis Tsiamas 외 arxiv

Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSONAR, a new family of omnilingual, cross-…

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

2025-01-10 · Fabian David Schmidt, Ivan Vulić, Goran Glavaš, David Ifeoluwa Adelani

While recent multilingual automatic speech recognition models claim to support thousands of languages, ASR for low-resource languages remains highly unreliable due to limited bimodal speech and text training data. Better…

Automatic Speech RecognitionClassificationintent-classificationIntent Classification+8

Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech

2022-10-27 · Takaaki Saeki, Heiga Zen, Zhehuai Chen, Nobuyuki Morioka 외

This paper proposes Virtuoso, a massively multilingual speech-text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech RecognitionSpeech Synthesis+3

Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages

2023-05-21 · Andrew Rouditchenko, Sameer Khurana, Samuel Thomas, Rogerio Feris 외

Recent models such as XLS-R and Whisper have made multilingual speech technologies more accessible by pre-training on audio from around 100 spoken languages each. However, there are thousands of spoken languages worldwid…

An Investigation of Hybrid architectures for Low Resource Multilingual Speech Recognition system in Indian context

2021-12-01 · ICON 2021 12 · Ganesh Mirishkar, Aditya Yadavalli, Anil Kumar Vuppala

India is a land of language diversity. There are approximately 2000 languages spoken around, and among which officially registered are 23. In those, there are very few with Automatic Speech Recognition (ASR) capability. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityLanguage Modeling+3