paper-with-me

홈 › Papers

WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

2024-06-09 · Linhan Ma, Dake Guo, Kun Song, Yuepeng Jiang, Shuai Wang, Liumeng Xue, Weiming Xu, Huan Zhao, BinBin Zhang, Lei Xie

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus derived from the open-sourced WenetSpeech dataset. Tailored for the text-to-speech tasks, we refined WenetSpeech by adjusting segment boundaries, enhancing the audio quality, and eliminating speaker mixing within each segment. Following a more accurate transcription process and quality-based data filtering process, the obtained WenetSpeech4TTS corpus contains $12,800$ hours of paired audio-text data. Furthermore, we have created subsets of varying sizes, categorized by segment quality scores to allow for TTS model training and fine-tuning. VALL-E and NaturalSpeech 2 systems are trained and fine-tuned on these subsets to validate the usability of WenetSpeech4TTS, establishing baselines on benchmark for fair comparison of TTS systems. The corpus and corresponding benchmarks are publicly available on huggingface.

📄 PDF Abstract BibTeX arXiv:2406.05763

Code (1)

dukGuo/valle-audiodec 공식 구현 pytorch

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition

2021-10-07 · BinBin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao 외

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours i…

Label Error DetectionOptical Character RecognitionOptical Character Recognition (OCR)speech-recognition+2

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

2025-09-22 · Yuhang Dai, Ziyu Zhang, Shuai Wang, Longhao Li 외 arxiv

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap…

TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a Speech Recognition Baseline

2022-06-27 · Chengfei Li, Shuhao Deng, Yaoping Wang, Guangjing Wang 외

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real on…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

VoiceBank-2023: A Multi-Speaker Mandarin Speech Corpus for Constructing Personalized TTS Systems for the Speech Impaired

2023-08-27 · Jia-Jyu Su, Pang-Chen Liao, Yen-Ting Lin, Wu-Hao Li 외

Services of personalized TTS systems for the Mandarin-speaking speech impaired are rarely mentioned. Taiwan started the VoiceBanking project in 2020, aiming to build a complete set of services to deliver personalized Man…

Introducing MELI: the Mandarin-English Language Interview Corpus

2026-03-27 · Suyuan Liu, Molly Babel arxiv

We introduce the Mandarin-English Language Interview (MELI) Corpus, an open-source resource of 29.8 hours of speech from 51 Mandarin-English bilingual speakers. MELI combines matched sessions in Mandarin and English with…