paper-with-me

홈 › Papers

ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis

2024-06-13 · Dehua Tao, Daxin Tan, Yu Ting Yeung, Xiao Chen, Tan Lee

Representing speech as discretized units has numerous benefits in supporting downstream spoken language processing tasks. However, the approach has been less explored in speech synthesis of tonal languages like Mandarin Chinese. Our preliminary experiments on Chinese speech synthesis reveal the issue of "tone shift", where a synthesized speech utterance contains correct base syllables but incorrect tones. To address the issue, we propose the ToneUnit framework, which leverages annotated data with tone labels as CTC supervision to learn tone-aware discrete speech units for Mandarin Chinese speech. Our findings indicate that the discrete units acquired through the TonUnit resolve the "tone shift" issue in synthesized Chinese speech and yield favorable results in English synthesis. Moreover, the experimental results suggest that finite scalar quantization enhances the effectiveness of ToneUnit. Notably, ToneUnit can work effectively even with minimal annotated data.

📄 PDF Abstract BibTeX arXiv:2406.08989

Code (0)

등록된 구현이 없습니다.

Tasks

QuantizationSpeech Synthesis

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

2026-07-05 · Offiong Bassey Edet, Emmanuel Oyo-Ita, Archibong Okon Archibong, David Effanga Bassey 외 arxiv

Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented en…

Speech Synthesis

Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages

2026-08-02 · Saierdaer Yusuyin, Nanling Jiang, Hao Huang, Zhijian Ou arxiv

Phoneme-based multilingual automatic speech recognition (ASR) can share acoustic evidence across languages more directly than language-specific subword modeling. When tonal and non-tonal languages are jointly trained, ho…

Language-Agnostic Analysis of Speech Depression Detection

2024-09-23 · Sona Binu, Jismi Jose, Fathima Shimna K V, Alino Luke Hans 외

The people with Major Depressive Disorder (MDD) exhibit the symptoms of tonal variations in their speech compared to the healthy counterparts. However, these tonal variations not only confine to the state of MDD but also…

Depression Detection

SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages

2026-01-14 · Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong 외 arxiv

Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech representations that are robust to nuisance variation, such as speaker gender, w…

Developing a Speech Recognition System for Recognizing Tonal Speech Signals Using a Convolutional Neural Network

2022-06-17 · 05-24 2022 6 · Sakshi Dua, Sethuraman Sambath Kumar, Yasser Albagory, Rajakumar Ramalingam 외

Deep learning-based machine learning models have shown significant results in speech recognition and numerous vision-related tasks. The performance of the present speech-to-text model relies upon the hyperparameters used…

speech-recognitionSpeech RecognitionSpeech-to-Text