paper-with-me

Papers

Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition

2020-10-26

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS synthesis on a TTS dataset without emotion labels. Specifically, our proposed method consists of a cross-domain speech emotion recognition (SER) model and an emotional TTS model. Firstly, we train the cross-domain SER model on both SER and TTS datasets. Then, we use emotion labels on the TTS dataset predicted by the trained SER model to build an auxiliary SER task and jointly train it with the TTS model. Experimental results show that our proposed method can generate speech with the specified emotional expressiveness and nearly no hindering on the speech quality.

📄 PDF Abstract BibTeX arXiv:2010.13350

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion RecognitionSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech

2025-05-15 · Jiaxuan Liu, ZhenHua Ling

Recent neural codec language models have made great progress in the field of text-to-speech (TTS), but controllable emotional TTS still faces many challenges. Traditional methods rely on predefined discrete emotion label…

Emotional Speech SynthesisLanguage ModelingLanguage ModellingSpeech Synthesis+2

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

2026-06-08 · Minghui Wu, Ganjun Liu, Zikun Fang, Ting Meng 외 arxiv

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained inte…

Speech Synthesis

ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models

2023-05-23 · Minki Kang, Wooseok Han, Sung Ju Hwang, Eunho Yang

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

2026-01-30 · Li Zhou, Hao Jiang, Junjie Li, Tianrui Wang 외 arxiv

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language…

Speech Synthesis

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

2024-09-16 · Xiaoxue Gao, Chen Zhang, Yiming Chen, Huayun Zhang 외

Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. Th…

Emotional Speech SynthesisIn-Context LearningInstruction FollowingSpeech Synthesis+2