paper-with-me

홈 › Papers

Speech Synthesis with Mixed Emotions

2022-08-11 · Kun Zhou, Berrak Sisman, Rajib Rana, B. W. Schuller, Haizhou Li

Emotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we seek to generate speech with a mixture of emotions at run-time. We propose a novel formulation that measures the relative difference between the speech samples of different emotions. We then incorporate our formulation into a sequence-to-sequence emotional text-to-speech framework. During the training, the framework does not only explicitly characterize emotion styles, but also explores the ordinal nature of emotions by quantifying the differences with other emotions. At run-time, we control the model to produce the desired emotion mixture by manually defining an emotion attribute vector. The objective and subjective evaluations have validated the effectiveness of the proposed framework. To our best knowledge, this research is the first study on modelling, synthesizing, and evaluating mixed emotions in speech.

📄 PDF Abstract BibTeX arXiv:2208.05890

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeEmotional Speech SynthesisSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

2025-06-03 · Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…

Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1

Mixed-EVC: Mixed Emotion Synthesis and Control in Voice Conversion

2022-10-25 · Kun Zhou, Berrak Sisman, Carlos Busso, Bin Ma 외

Emotional voice conversion (EVC) traditionally targets the transformation of spoken utterances from one emotional state to another, with previous research mainly focusing on discrete emotion categories. This paper depart…

AttributeVoice Conversion

SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate

2022-07-13 · Nabarun Goswami, Tatsuya Harada

The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylistic factors like gender, age, accent, em…

Speech Separationtext-to-speechText to Speech

Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations

2022-11-11 · Yoori Oh, Juheon Lee, Yoseob Han, Kyogu Lee

Recent text-to-speech models have reached the level of generating natural speech similar to what humans say. But there still have limitations in terms of expressiveness. The existing emotional speech synthesis models hav…

Emotional Speech SynthesisSpeech Synthesistext-to-speechText to Speech

SyntAct: A Synthesized Database of Basic Emotions

2022-06-01 · DCLRL (LREC) 2022 6 · Felix Burkhardt, Florian Eyben, Björn Schuller

Speech emotion recognition is in the focus of research since several decades and has many applications. One problem is sparse data for supervised learning. One way to tackle this problem is the synthesis of data with emo…

Emotion RecognitionSpeech Emotion RecognitionSpeech Synthesis