paper-with-me

Papers

Naturalistic Head Motion Generation from Speech

2022-10-26 · Trisha Mittal, Zakaria Aldeneh, Masha Fedzechkina, Anurag Ranjan, Barry-John Theobald

Synthesizing natural head motion to accompany speech for an embodied conversational agent is necessary for providing a rich interactive experience. Most prior works assess the quality of generated head motion by comparing them against a single ground-truth using an objective metric. Yet there are many plausible head motion sequences to accompany a speech utterance. In this work, we study the variation in the perceptual quality of head motions sampled from a generative model. We show that, despite providing more diverse head motions, the generative model produces motions with varying degrees of perceptual quality. We finally show that objective metrics commonly used in previous research do not accurately reflect the perceptual quality of generated head motions. These results open an interesting avenue for future work to investigate better objective metrics that correlate with human perception of quality.

📄 PDF Abstract BibTeX arXiv:2210.14800

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Generation

Similar Papers 제목 키워드 기반

ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge

2025-05-23 · Soumya Dutta, Smruthi Balaji, Varada R, Viveka Salinamakki 외

Speech emotion recognition (SER) in naturalistic settings remains a challenge due to the intrinsic variability, diverse recording conditions, and class imbalance. As participants in the Interspeech Naturalistic SER Chall…

Emotion RecognitionSpeech Emotion Recognition

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

2025-06-11 · Georgios Chatzichristodoulou, Despoina Kosmopoulou, Antonios Kritikos, Anastasia Poulopoulou 외

SER is a challenging task due to the subjective nature of human emotions and their uneven representation under naturalistic conditions. We propose MEDUSA, a multimodal framework with a four-stage training pipeline, which…

Emotion RecognitionSpeech Emotion Recognition

Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction

2025-06-12 · Thanathai Lertpetchpun, Tiantian Feng, Dani Byrd, Shrikanth Narayanan

Speech emotion recognition (SER) in naturalistic conditions presents a significant challenge for the speech processing community. Challenges include disagreement in labeling among annotators and imbalanced data distribut…

AttributeEmotion RecognitionMulti-Task LearningSpeech Emotion Recognition+1

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

2025-06-02 · Alef Iury Siqueira Ferreira, Lucas Rafael Gris, Alexandre Ferro Filho, Lucas Ólives 외

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the IN…

Audio TaggingEmotion RecognitionGraph AttentionQuantization+1

SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation

2024-05-12 · Changpeng Cai, Guinan Guo, Jiao Li, Junhao Su 외

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. Whil…

DisentanglementFace GenerationTalking Face GenerationTalking Head Generation