paper-with-me

홈 › Papers

OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

2026-06-08 · David Guzmán, Luel Hagos Beyene, Jesujoba Oluwadara Alabi, Yejin Jeon, Dietrich Klakow, David Ifeoluwa Adelani arxiv

Recent advances in neural text-to-speech (TTS) and multilingual speech generation have substantially improved synthetic speech quality, yet these gains remain unevenly distributed across the world's languages. Existing models are still dominated by a small set of high-resource languages, while many studies of low-resource TTS are simulated on artificially downsampled high-resource corpora that do not reflect the orthographic variation and limited phonetic coverage encountered in genuinely underrepresented settings. As such, we introduce OpenBibleTTS, which is a large-scale benchmark for low-resource speech synthesis spanning 37 underrepresented languages. Moreover, a systematic comparison of various TTS architectures and large-scale speech generation models is conducted across in-domain Biblical text and out-of-domain material. Results show that no single system dominates across languages and metrics: Gemini-TTS achieves the highest listener ratings on most evaluated languages, but monolingual EveryVoice models trained on OpenBibleTTS remain strongest for intelligibility and are preferred in several African languages, while open from-scratch systems degrade sharply on out-of-domain text, revealing a persistent gap between broad multilingual coverage and reliable synthesis quality in underserved linguistic communities. We complement automatic evaluation with subjective human judgments, and open-source all processed datasets, alignments, and trained models to support future low-resource TTS research.

📄 PDF Abstract BibTeX arXiv:2606.09553

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

2025-06-01 · Taesoo Kim, Jong Hwan Ko

Recent advances in speech-enabled language models have shown promising results in building intelligent voice assistants. However, most existing approaches rely on large-scale paired speech-text data and extensive computa…

ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams

2024-10-23 · Srija Anand, Praveen Srinivasa Varadhan, Mehak Singal, Mitesh M. Khapra

Recent advancements in Text-to-Speech (TTS) technology have led to natural-sounding speech for English, primarily due to the availability of large-scale, high-quality web data. However, many other languages lack access t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingKnowledge Distillation+5

OSUM: Advancing Open Speech Understanding Models with Limited Resources in Academia

2025-01-23 · Xuelong Geng, Kun Wei, Qijie Shao, Shuiyun Liu 외

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions. Howeve…

Emotion RecognitionEvent DetectionGender ClassificationSpeech Emotion Recognition+3

Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information

2012-05-01 · LREC 2012 5 · Chi-Hsin Yu, Yi-jie Tang, Hsin-Hsi Chen

Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variat…

Information RetrievalLanguage Modelling

VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

2026-06-23 · Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho, Phuong Tuan Dat 외 arxiv

Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial…

Speaker Recognition