paper-with-me

홈 › Papers

On the Emotion Understanding of Synthesized Speech

2026-03-17 · Yuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang, Bei Li, Xiaoqian Liu, Chenglong Wang, Jianjin Wang, Bingsen Zhou, Bingyu Liu, Jingbo Zhu, Zhengtao Yu, Tong Xiao arxiv

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding results a plausible reward or evaluation metric for assessing emotional expressiveness in speech synthesis. In this work, we critically examine this assumption by systematically evaluating Speech Emotion Recognition (SER) on synthesized speech across datasets, discriminative and generative SER models, and diverse synthesis models. We find that current SER models can not generalize to synthesized speech, largely because speech token prediction during synthesis induces a representation mismatch between synthesized and human speech. Moreover, generative Speech Language Models (SLMs) tend to infer emotion from textual semantics while ignoring paralinguistic cues. Overall, our findings suggest that existing SER models often exploit non-robust shortcuts rather than capturing fundamental features, and paralinguistic understanding in SLMs remains challenging.

📄 PDF Abstract BibTeX arXiv:2603.16483

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion RecognitionSpeech Synthesis

Similar Papers 제목 키워드 기반

SyntAct: A Synthesized Database of Basic Emotions

2022-06-01 · DCLRL (LREC) 2022 6 · Felix Burkhardt, Florian Eyben, Björn Schuller

Speech emotion recognition is in the focus of research since several decades and has many applications. One problem is sparse data for supervised learning. One way to tackle this problem is the synthesis of data with emo…

Emotion RecognitionSpeech Emotion RecognitionSpeech Synthesis

Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding

2025-02-26 · Tianyun Liu

Traditional text-to-speech (TTS) methods primarily focus on establishing a mapping between phonemes and mel-spectrograms. However, during the phoneme encoding stage, there is often a lack of real mel-spectrogram auxiliar…

text-to-speechText to Speech

Emotional Prosody Control for Speech Generation

2021-11-07 · Sarath Sivaprasad, Saiteja Kosgi, Vineet Gandhi

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average varia…

text-to-speechText to Speech

Multi-speaker Emotional Text-to-speech Synthesizer

2021-12-07 · Sungjae Cho, Soo-Young Lee

We present a methodology to train our multi-speaker emotional text-to-speech synthesizer that can express speech for 10 speakers' 7 different emotions. All silences from audio samples are removed prior to learning. This …

Alltext-to-speechText to Speech

EmoSpeech: Guiding FastSpeech2 Towards Emotional Text to Speech

2023-06-28 · Daria Diatlova, Vitaly Shutov

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the st…

Emotion RecognitionSpeech Synthesistext-to-speechText to Speech