paper-with-me

홈 › Papers

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

2025-09-30 · Kang Yang, Yifan Liang, Fangkun Liu, Zhenping Xie, Chengshi Zheng arxiv

Lip-to-speech (L2S) synthesis for Mandarin is a significant challenge, hindered by complex viseme-to-phoneme mappings and the critical role of lexical tones in intelligibility. To address this issue, we propose Lexical Tone-Aware Lip-to-Speech (LTA-L2S). To tackle viseme-to-phoneme complexity, our model adapts an English pre-trained audio-visual self-supervised learning (SSL) model via a cross-lingual transfer learning strategy. This strategy not only transfers universal knowledge learned from extensive English data to the Mandarin domain but also circumvents the prohibitive cost of training such a model from scratch. To specifically model lexical tones and enhance intelligibility, we further employ a flow-matching model to generate the F0 contour. This generation process is guided by ASR-fine-tuned SSL speech units, which contain crucial suprasegmental information. The overall speech quality is then elevated through a two-stage training paradigm, where a flow-matching postnet refines the coarse spectrogram from the first stage. Extensive experiments demonstrate that LTA-L2S significantly outperforms existing methods in both speech intelligibility and tonal accuracy.

📄 PDF Abstract BibTeX arXiv:2509.25670

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningCross-Lingual TransferSpeech Synthesis

Similar Papers 제목 키워드 기반

ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis

2024-06-13 · Dehua Tao, Daxin Tan, Yu Ting Yeung, Xiao Chen 외

Representing speech as discretized units has numerous benefits in supporting downstream spoken language processing tasks. However, the approach has been less explored in speech synthesis of tonal languages like Mandarin …

QuantizationSpeech Synthesis

Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yorùbá

2026-04-08 · Opeyemi Osakuade, Simon King arxiv

Discrete speech units (DSUs) are derived by quantising representations from models trained using self-supervised learning (SSL). They are a popular representation for a wide variety of spoken language tasks, including th…

Self-Supervised LearningRepresentation Learning

Phonetic and semantic analyses of spoken corpora of Beijing and Taiwan Mandarin indicate that the neutral tone is a lexical tone

2026-06-24 · Yuxin Lu, Zhexuan Li, R. Harald Baayen arxiv

The neutral, or floating, tone of Mandarin Chinese is a tone with an enigmatic set of properties. It has been described as a reduced tone, or as a tone that sometimes is lexically fixed but that can also be toneless. In …

A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models

2024-08-24 · Antón de la Fuente, Dan Jurafsky

This study asks how self-supervised speech models represent suprasegmental categories like Mandarin lexical tone, English lexical stress, and English phrasal accents. Through a series of probing tasks, we make layer-wise…

Specificity

SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages

2026-01-14 · Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong 외 arxiv

Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech representations that are robust to nuisance variation, such as speaker gender, w…