paper-with-me

홈 › Papers

MLS: A Large-Scale Multilingual Dataset for Speech Research

2020-12-07 · Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, Ronan Collobert

This paper introduces Multilingual LibriSpeech (MLS) dataset, a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages, including about 44.5K hours of English and a total of about 6K hours for other languages. Additionally, we provide Language Models (LM) and baseline Automatic Speech Recognition (ASR) models and for all the languages in our dataset. We believe such a large transcribed dataset will open new avenues in ASR and Text-To-Speech (TTS) research. The dataset will be made freely available for anyone at http://www.openslr.org.

📄 PDF Abstract BibTeX arXiv:2012.03411

Code (1)

facebookresearch/wav2letter/tree/master/recipes/mls 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

CoVoST 2 and Massively Multilingual Speech-to-Text Translation

2020-07-20 · Changhan Wang, Anne Wu, Juan Pino

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to f…

Machine Translationspeech-recognitionSpeech RecognitionSpeech-to-Text+2

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition

OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics

2025-09-04 · Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han 외 arxiv

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and ot…

Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages

2023-06-07 · Claytone Sikasote, Kalinda Siaminwe, Stanly Mwape, Bangiwe Zulu 외

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours…

Cross-Lingual Transferspeech-recognitionSpeech RecognitionTransfer Learning

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

2025-01-27 · Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li 외

Recent advancements in speech generation have been driven by the large-scale training datasets. However, current models fall short of capturing the spontaneity and variability inherent in real-world human speech, due to …