paper-with-me

홈 › Papers

Emotion-Aware Prosodic Phrasing for Expressive Text-to-Speech

2023-09-21 · Rui Liu, Bin Liu, Haizhou Li

Prosodic phrasing is crucial to the naturalness and intelligibility of end-to-end Text-to-Speech (TTS). There exist both linguistic and emotional prosody in natural speech. As the study of prosodic phrasing has been linguistically motivated, prosodic phrasing for expressive emotion rendering has not been well studied. In this paper, we propose an emotion-aware prosodic phrasing model, termed \textit{EmoPP}, to mine the emotional cues of utterance accurately and predict appropriate phrase breaks. We first conduct objective observations on the ESD dataset to validate the strong correlation between emotion and prosodic phrasing. Then the objective and subjective evaluations show that the EmoPP outperforms all baselines and achieves remarkable performance in terms of emotion expressiveness. The audio samples and the code are available at \url{https://github.com/AI-S2-Lab/EmoPP}.

📄 PDF Abstract BibTeX arXiv:2309.11724

Code (1)

ai-s2-lab/emopp 공식 구현

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech

2025-09-25 · Aurosweta Mahapatra, Ismail Rasim Ulgen, Berrak Sisman arxiv

Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans…

Self-Supervised LearningMulti-Task LearningSpoof Detection

ADEPT: A Dataset for Evaluating Prosody Transfer

2021-06-15 · Alexandra Torresquintero, Tian Huey Teh, Christopher G. R. Wallis, Marlene Staib 외

Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a reference speech sample. There have been consid…

text-to-speechText to Speech

A Discourse-level Multi-scale Prosodic Model for Fine-grained Emotion Analysis

2023-09-21 · Xianhao Wei, Jia Jia, Xiang Li, Zhiyong Wu 외

This paper explores predicting suitable prosodic features for fine-grained emotion analysis from the discourse-level text. To obtain fine-grained emotional prosodic features as predictive values for our model, we extract…

Emotion RecognitionSpeech SynthesisStyle Transfer

RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis

2025-10-16 · Qing Yang, Zhenghao Liu, Yangfan Du, Pengcheng Huang 외 arxiv

Recent advances in Text-To-Speech (TTS) synthesis have achieved near-human speech quality in neutral speaking styles. However, most existing approaches either depend on costly emotion annotations or optimize surrogate ob…

Reinforcement LearningSpeech RecognitionSpeech Synthesis

CNN Encoding of Acoustic Parameters for Prominence Detection

2021-04-12 · Kamini Sabu, Mithilesh Vaidya, Preeti Rao

Expressive reading, considered the defining attribute of oral reading fluency, comprises the prosodic realization of phrasing and prominence. In the context of evaluating oral reading, it helps to establish the speaker's…

Attribute