paper-with-me

홈 › Papers

Human Feedback Driven Dynamic Speech Emotion Recognition

2025-08-18 · Ilya Fedorov, Dmitry Korobchenko arxiv

This work proposes to explore a new area of dynamic speech emotion recognition. Unlike traditional methods, we assume that each audio track is associated with a sequence of emotions active at different moments in time. The study particularly focuses on the animation of emotional 3D avatars. We propose a multi-stage method that includes the training of a classical speech emotion recognition model, synthetic generation of emotional sequences, and further model improvement based on human feedback. Additionally, we introduce a novel approach to modeling emotional mixtures based on the Dirichlet distribution. The models are evaluated based on ground-truth emotions extracted from a dataset of 3D facial animations. We compare our models against the sliding window approach. Our experimental results show the effectiveness of Dirichlet-based approach in modeling emotional mixtures. Incorporating human feedback further improves the model quality while providing a simplified annotation procedure.

📄 PDF Abstract BibTeX arXiv:2508.14920

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion Recognition

Similar Papers 제목 키워드 기반

RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis

2025-10-16 · Qing Yang, Zhenghao Liu, Yangfan Du, Pengcheng Huang 외 arxiv

Recent advances in Text-To-Speech (TTS) synthesis have achieved near-human speech quality in neutral speaking styles. However, most existing approaches either depend on costly emotion annotations or optimize surrogate ob…

Reinforcement LearningSpeech RecognitionSpeech Synthesis

Audio-Driven Emotional Video Portraits

2021-04-15 · CVPR 2021 1 · Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu 외

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important feat…

DisentanglementFace Generation

Speech Driven Talking Face Generation from a Single Image and an Emotion Condition

2020-08-08 · Sefik Emre Eskimez, You Zhang, Zhiyao Duan

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifical…

Emotion RecognitionFace GenerationTalking Face Generation

AMUSE: Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2024-06-01 · IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) 2024 6 · Chhatre K., Danecek R., Athanasiou N., Becherini G. 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm

Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2023-12-07 · CVPR 2024 1 · Kiran Chhatre, Radek Daněček, Nikos Athanasiou, Giorgio Becherini 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm