paper-with-me

홈 › Papers

No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS

2025-09-23 · Seungyoun Shin, Dongha Ahn, Jiwoo Kim, Sungwook Jeon arxiv

Recent work reports gains in neural text-to-speech (TTS) with Group Relative Policy Optimization (GRPO). However, in the absence of a verifiable reward for \textit{prosody}, GRPO trained on transcription-oriented signals (CER/NLL) lowers error rates yet collapses prosody into monotone, unnatural speech; adding speaker-similarity further destabilizes training and degrades CER. We address this with an \textit{iterative Direct Preference Optimization (DPO)} scheme that uses only a few hundred human-labeled preference pairs per round to directly optimize prosodic naturalness while regularizing to the current model. On \textbf{KoCC-TTS}, a curated dataset of authentic Korean call center interactions capturing task-oriented dialogues, our method attains the highest human preference (ELO) with competitive CER, outperforming GRPO and strong commercial baselines. These results suggest that when prosody cannot be rewarded automatically, \textit{human preference optimization} offers a practical and data-efficient path to natural and robust TTS. The demo page is available at \href{https://tts.ch.dev}

📄 PDF Abstract BibTeX arXiv:2509.18531

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Speech BERT Embedding For Improving Prosody in Neural TTS

2021-06-08 · Liping Chen, Yan Deng, Xi Wang, Frank K. Soong 외

This paper presents a speech BERT model to extract embedded prosody information in speech segments for improving the prosody of synthesized speech in neural text-to-speech (TTS). As a pre-trained model, it can learn pros…

Decodertext-to-speechText to Speech

Global Rhythm Style Transfer Without Text Transcriptions

2021-06-16 · Kaizhi Qian, Yang Zhang, Shiyu Chang, JinJun Xiong 외

Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody information. Two major components of pro…

Representation LearningRhythmStyle Transfer

eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer

2023-06-20 · Ammar Abbas, Sri Karlapati, Bastian Schnell, Penny Karanasou 외

We present eCat, a novel end-to-end multispeaker model capable of: a) generating long-context speech with expressive and contextually appropriate prosody, and b) performing fine-grained prosody transfer between any pair …

Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale

2025-11-26 · Yicheng Zhong, Peiji Yang, Zhisheng Wang arxiv

Recent advances in Large Language Models (LLMs) have transformed text-to-speech (TTS) synthesis, inspiring autoregressive frameworks that represent speech as sequences of discrete codec tokens. Among them, single-codeboo…

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs

2025-08-12 · Eray Eren, Qingju Liu, Hyeongwoo Kim, Pablo Garrido 외 arxiv

Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic features such as F0 and energy and can be used …