paper-with-me

홈 › Papers

Improving French Synthetic Speech Quality via SSML Prosody Control

2025-08-24 · Nassima Ould Ouali, Awais Hussain Sani, Ruben Bueno, Jonah Dauvet, Tim Luka Horstmann, Eric Moulines arxiv

Despite recent advances, synthetic voices often lack expressiveness due to limited prosody control in commercial text-to-speech (TTS) systems. We introduce the first end-to-end pipeline that inserts Speech Synthesis Markup Language (SSML) tags into French text to control pitch, speaking rate, volume, and pause duration. We employ a cascaded architecture with two QLoRA-fine-tuned Qwen 2.5-7B models: one predicts phrase-break positions and the other performs regression on prosodic targets, generating commercial TTS-compatible SSML markup. Evaluated on a 14-hour French podcast corpus, our method achieves 99.2% F1 for break placement and reduces mean absolute error on pitch, rate, and volume by 25-40% compared with prompting-only large language models (LLMs) and a BiLSTM baseline. In perceptual evaluation involving 18 participants across over 9 hours of synthesized audio, SSML-enhanced speech generated by our pipeline significantly improves naturalness, with the mean opinion score increasing from 3.20 to 3.87 (p < 0.005). Additionally, 15 of 18 listeners preferred our enhanced synthesis. These results demonstrate substantial progress in bridging the expressiveness gap between synthetic and natural French speech. Our code is publicly available at https://github.com/hi-paris/Prosody-Control-French-TTS.

📄 PDF Abstract BibTeX arXiv:2508.17494

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution

2025-10-27 · Dharma Teja Donepudi arxiv

Intra-sentence multilingual speech synthesis (code-switching TTS) remains a major challenge due to abrupt language shifts, varied scripts, and mismatched prosody between languages. Conventional TTS systems are typically …

Language IdentificationSpeech Synthesis

La variation prosodique dialectale en fran\ccais. Donn\'ees et hypoth\`eses (Speech Prosody of Dialectal French: Data and Hypotheses) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Mathieu Avanzi, Nicolas Obin, Guri Bordal, Alice Bardiaux

Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages

2026-08-28 · Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman 외 arxiv

Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST…

Speech-to-Speech Translation

La structuration prosodique et les relations syntaxe/ prosodie dans le discours politique (Prosodic Structuring and the Syntax-Prosody Relationship in Political Speech) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Ingo Feldhausen, Elisabeth Delais-Roussarie
Chunking

ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech

2022-02-16 · Yi Ren, Ming Lei, Zhiying Huang, Shiliang Zhang 외

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling w…

text-to-speechText to Speech