paper-with-me

Papers

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation

2025-09-22 · Yutong Liu, Ziyue Zhang, Ban Ma-bao, Renzeng Duojie, Yuqing Cai, Yongbin Yu, Xiangxiang Wang, Fan Gao, Cheng Huang, Nyima Tashi arxiv

Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (Ü-Tsang, Amdo, and Kham), limiting progress in speech modeling. To address this issue, we propose TMD-TTS, a unified Tibetan multi-dialect text-to-speech (TTS) framework that synthesizes parallel dialectal speech from explicit dialect labels. Our method features a dialect fusion module and a Dialect-Specialized Dynamic Routing Network (DSDR-Net) to capture fine-grained acoustic and linguistic variations across dialects. Extensive objective and subjective evaluations demonstrate that TMD-TTS significantly outperforms baselines in dialectal expressiveness. We further validate the quality and utility of the synthesized speech through a challenging Speech-to-Speech Dialect Conversion (S2SDC) task.

📄 PDF Abstract BibTeX arXiv:2509.18060

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation

2026-05-04 · Jiaxu He, Chao Wang, Jie Lian, Yuqing Cai 외 arxiv

Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between written text and spoken pronunciation. To address these issues, this work…

Speech Synthesis

FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation

2025-05-20 · Yutong Liu, Ziyue Zhang, Ban Ma-bao, Yuqing Cai 외

Tibetan is a low-resource language with minimal parallel speech corpora spanning its three major dialects-\"U-Tsang, Amdo, and Kham-limiting progress in speech modeling. To address this issue, we propose FMSD-TTS, a few-…

Dataset GenerationSpeech Synthesistext-to-speechText to Speech+1

Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe

2025-08-03 · Tiantian Feng, Kevin Huang, Anfeng Xu, Xuan Shi 외 arxiv

We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language v…

Speech Recognition

Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges

2025-10-22 · Cheng Huang, Nyima Tashi, Fan Gao, Yutong Liu 외 arxiv

Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research. Despite increasing interest in developin…

Cross-Lingual TransferMachine TranslationSpeech Recognition

Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

2026-01-20 · Yushen Chen, Junzhe Liu, Yujie Tu, Zhikang Niu 외 arxiv

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absen…

Speech Synthesis