paper-with-me

홈 › Papers

Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning

2022-06-15 · Rui Liu, Berrak Sisman, Björn Schuller, Guanglai Gao, Haizhou Li

Emotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was proposed to predict emotion strength for emotional speech corpus. However, the trained ranking function doesn't generalize to new domains, which limits the scope of applications, especially for out-of-domain or unseen speech. In this paper, we propose a data-driven deep learning model, i.e. StrengthNet, to improve the generalization of emotion strength assessment for seen and unseen speech. This is achieved by the fusion of emotional data from various domains. We follow a multi-task learning network architecture that includes an acoustic encoder, a strength predictor, and an auxiliary emotion predictor. Experiments show that the predicted emotion strength of the proposed StrengthNet is highly correlated with ground truth scores for both seen and unseen speech. We release the source codes at: https://github.com/ttslr/StrengthNet.

📄 PDF Abstract BibTeX arXiv:2206.07229

Code (1)

ttslr/strengthnet 공식 구현 tf

Tasks

AttributeEmotion ClassificationMulti-Task Learningtext-to-speechText to SpeechVoice Conversion

Similar Papers 제목 키워드 기반

StrengthNet: Deep Learning-based Emotion Strength Assessment for Emotional Speech Synthesis

2021-10-07 · Rui Liu, Berrak Sisman, Haizhou Li

Recently, emotional speech synthesis has achieved remarkable performance. The emotion strength of synthesized speech can be controlled flexibly using a strength descriptor, which is obtained by an emotion attribute ranki…

AttributeData AugmentationEmotional Speech SynthesisMulti-Task Learning+1

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

2025-06-03 · Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…

Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1

Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversation

2025-08-27 · Kun Peng, Cong Cao, Hao Peng, Guanlin Wu 외 arxiv

Current Emotion Recognition in Conversation (ERC) research follows a closed-domain assumption. However, there is no clear consensus on emotion classification in psychology, which presents a challenge for models when it c…

Emotion Recognition in ConversationEmotion Classification

Nonparallel Emotional Voice Conversion For Unseen Speaker-Emotion Pairs Using Dual Domain Adversarial Network & Virtual Domain Pairing

2023-02-21 · Nirmesh Shah, Mayank Kumar Singh, Naoya Takahashi, Naoyuki Onoe

Primary goal of an emotional voice conversion (EVC) system is to convert the emotion of a given speech signal from one style to another style without modifying the linguistic content of the signal. Most of the state-of-t…

Voice Conversion

Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models

2026-03-18 · Xiutian Zhao, Ismail Rasim Ulgen, Philipp Koehn, Björn Schuller 외 arxiv

Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic fidelity through refusals, hallucinatio…