paper-with-me

홈 › Papers

XTREME-S: Evaluating Cross-lingual Speech Representations

2022-03-21 · Alexis Conneau, Ankur Bapna, Yu Zhang, Min Ma, Patrick von Platen, Anton Lozhkov, Colin Cherry, Ye Jia, Clara Rivera, Mihir Kale, Daan van Esch, Vera Axelrod, Simran Khanuja, Jonathan H. Clark, Orhan Firat, Michael Auli, Sebastian Ruder, Jason Riesa, Melvin Johnson

We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval. Covering 102 languages from 10+ language families, 3 different domains and 4 task families, XTREME-S aims to simplify multilingual speech representation evaluation, as well as catalyze research in "universal" speech representation learning. This paper describes the new benchmark and establishes the first speech-only and speech-text baselines using XLS-R and mSLAM on all downstream tasks. We motivate the design choices and detail how to use the benchmark. Datasets and fine-tuning scripts are made easily accessible at https://hf.co/datasets/google/xtreme_s.

📄 PDF Abstract BibTeX arXiv:2203.10752

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningRetrievalspeech-recognitionSpeech RecognitionSpeech Representation LearningSpeech-to-TextSpeech-to-Text TranslationTranslation

Similar Papers 제목 키워드 기반

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

2020-03-24 · Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 외

Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited …

Cross-Lingual TransferRetrievalSentenceSentence Retrieval

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation

2020-01-01 · ICML 2020 1 · Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 외

Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited …

Cross-Lingual TransferRetrievalSentenceSentence Retrieval+1

Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder

2026-03-09 · Maryem Bouziane, Salima Mdhaffar, Yannick Estève arxiv

Speech foundation models trained with self-supervised learning produce generic speech representations that support a wide range of speech processing tasks. When further adapted with supervised learning, these models can …

Self-Supervised LearningSpeaker Recognition

Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease

2026-03-23 · Abner Hernandez, Eunjung Yeo, Kwanghee Choi, Chin-Jou Li 외 arxiv

The limited availability of dysarthric speech data makes cross-lingual detection an important but challenging problem. A key difficulty is that speech representations often encode language-dependent structure that can co…

CLSRIL-23: Cross Lingual Speech Representations for Indic Languages

2021-07-15 · Anirudh Gupta, Harveen Singh Chadha, Priyanshi Shah, Neeraj Chhimwal 외

We present a CLSRIL-23, a self supervised learning based audio pre-trained model which learns cross lingual speech representations from raw audio across 23 Indic languages. It is built on top of wav2vec 2.0 which is solv…

Self-Supervised Learningspeech-recognitionSpeech Recognition