paper-with-me

홈 › Papers

Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models

2025-06-01 · Kyowoon Lee, Artyom Stitsyuk, Gunu Jho, Inchul Hwang, Jaesik Choi

Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation often depend on specialized modules or additional training, limiting their capacity for post-hoc adjustments. Similarly, traditional mispronunciation correction relies on grapheme-to-phoneme dictionaries, making it less practical in low-resource settings. We introduce Counterfactual Activation Editing, a model-agnostic method that manipulates internal representations in a pre-trained TTS model to achieve post-hoc control of prosody and pronunciation. Experimental results show that our method effectively adjusts prosodic features and corrects mispronunciations while preserving synthesis quality. This opens the door to inference-time refinement of TTS outputs without retraining, bridging the gap between pre-trained TTS models and editable speech synthesis.

📄 PDF Abstract BibTeX arXiv:2506.00832

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Speech Editing -- a Summary

2024-07-24 · Tobias Kässmann, Yining Liu, Danni Liu

With the rise of video production and social media, speech editing has become crucial for creators to address issues like mispronunciations, missing words, or stuttering in audio recordings. This paper explores text-base…

Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2

2026-03-12 · Suvendu Sekhar Mohanty arxiv

We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual tra…

Speech Synthesis

Context-Aware Prosody Correction for Text-Based Speech Editing

2021-02-16 · Max Morrison, Lucas Rencker, Zeyu Jin, Nicholas J. Bryan 외

Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is tha…

Denoising

Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation

2025-08-22 · Xueyao Zhang, Junan Zhang, Yuancheng Wang, Chaoren Wang 외 arxiv

Controllable human voice generation, particularly for expressive domains like singing, remains a significant challenge. This paper introduces Vevo2, a unified framework for controllable speech and singing voice generatio…

MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering

2026-03-07 · Trong-Thang Pham, Loc Nguyen, Anh Nguyen, Hien Nguyen 외 arxiv

Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training data. Re-prompting rerolls the entire generation trajectory, altering anatomy, te…

Data Augmentation