paper-with-me

홈 › Papers

Fine-grained Emotional Control of Text-To-Speech: Learning To Rank Inter- And Intra-Class Emotion Intensities

2023-03-02 · Shijun Wang, Jón Guðnason, Damian Borth

State-of-the-art Text-To-Speech (TTS) models are capable of producing high-quality speech. The generated speech, however, is usually neutral in emotional expression, whereas very often one would want fine-grained emotional control of words or phonemes. Although still challenging, the first TTS models have been recently proposed that are able to control voice by manually assigning emotion intensity. Unfortunately, due to the neglect of intra-class distance, the intensity differences are often unrecognizable. In this paper, we propose a fine-grained controllable emotional TTS, that considers both inter- and intra-class distances and be able to synthesize speech with recognizable intensity difference. Our subjective and objective experiments demonstrate that our model exceeds two state-of-the-art controllable TTS models for controllability, emotion expressiveness and naturalness.

📄 PDF Abstract BibTeX arXiv:2303.01508

Code (0)

등록된 구현이 없습니다.

Tasks

Learning-To-Ranktext-to-speechText to Speech

Similar Papers 제목 키워드 기반

EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering

2025-08-05 · Tianxin Xie, Shan Yang, Chenxing Li, Dong Yu 외 arxiv

Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emo…

Continuous Control

QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis

2023-03-14 · Haobin Tang, xulong Zhang, Jianzong Wang, Ning Cheng 외

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and cont…

Emotional Speech SynthesisSentenceSpeech Synthesistext-to-speech+1

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

2025-09-20 · Sirui Wang, Andong Chen, Tiejun Zhao arxiv

Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predefined labels, reference audio, or natural…

Speech Synthesis

Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement

2025-10-02 · Jianing Yang, Sheng Li, Takahiro Shinozaki, Yuki Saito 외 arxiv

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, w…

Style Transfer

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

2026-06-08 · Minghui Wu, Ganjun Liu, Zikun Fang, Ting Meng 외 arxiv

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained inte…

Speech Synthesis