paper-with-me

Papers

Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation

2024-08-01 · Xinhan Di, Zihao Chen, Yunming Liang, Junjie Zheng, Yihua Wang, Chaofan Ding

Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we propose Bailing-TTS, a family of large-scale TTS models capable of generating high-quality Chinese dialectal speech. Bailing-TTS serves as a foundation model for Chinese dialectal speech generation. First, continual semi-supervised learning is proposed to facilitate the alignment of text tokens and speech tokens. Second, the Chinese dialectal representation learning is developed using a specific transformer architecture and multi-stage training processes. With the proposed design of novel network architecture and corresponding strategy, Bailing-TTS is able to generate Chinese dialectal speech from text effectively and efficiently. Experiments demonstrate that Bailing-TTS generates Chinese dialectal speech towards human-like spontaneous representation. Readers are encouraged to listen to demos at \url{https://c9412600.github.io/bltts_tech_report/index.html}.

📄 PDF Abstract BibTeX arXiv:2408.00284

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis

2025-12-21 · Pengchao Feng, Yao Xiao, Ziyang Ma, Zhikang Niu 외 arxiv

Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted toward enhancing the expressiveness of ge…

Speech Synthesis

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

2025-09-22 · Yuhang Dai, Ziyu Zhang, Shuai Wang, Longhao Li 외 arxiv

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap…

Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

2026-01-20 · Yushen Chen, Junzhe Liu, Yujie Tu, Zhikang Niu 외 arxiv

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absen…

Speech Synthesis

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

2026-02-17 · Ahmed Khaled Khamis, Hesham Ali arxiv

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic …

Synthetic Data GenerationSpeaker DiarizationSpeech SynthesisText to Speech

Towards Zero-Shot Text-To-Speech for Arabic Dialects

2024-06-24 · Khai Duy Doan, Abdul Waheed, Muhammad Abdul-Mageed

Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native s…

Dialect IdentificationSpeech Synthesistext-to-speechText to Speech