paper-with-me

홈 › Papers

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

2024-07-07 · Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, Yuancheng Wang, Kai Chen, Pengyuan Zhang, Zhizheng Wu

Recent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/.

📄 PDF Abstract BibTeX arXiv:2407.05361

Code (1)

open-mmlab/Amphion/blob/main/preprocessors/Emilia/README.md 공식 구현 pytorch

Tasks

Text to Speech

Similar Papers 제목 키워드 기반

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

2025-01-27 · Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li 외

Recent advancements in speech generation have been driven by the large-scale training datasets. However, current models fall short of capturing the spontaneity and variability inherent in real-world human speech, due to …

SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models

2025-09-15 · Karan Dua, Puneet Mittal, Ranjeet Gupta, Hitesh Laxmichand Patel arxiv

High-quality Text-to-Speech (TTS) model training requires extensive and diverse text and speech data. It is challenging to procure such data from real sources due to issues of domain specificity, licensing, and scalabili…

Text to Speech

OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics

2025-09-04 · Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han 외 arxiv

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and ot…

MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models

2024-11-27 · Thai-Binh Nguyen, Alexander Waibel

Speaker-attributed automatic speech recognition (SA-ASR) aims to transcribe speech while assigning transcripts to the corresponding speakers accurately. Existing methods often rely on complex modular systems or require e…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

CoVoST 2 and Massively Multilingual Speech-to-Text Translation

2020-07-20 · Changhan Wang, Anne Wu, Juan Pino

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to f…

Machine Translationspeech-recognitionSpeech RecognitionSpeech-to-Text+2