paper-with-me

Papers

EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation

2023-05-30 · Xingqun Qi, Chen Liu, Lincheng Li, Jie Hou, Haoran Xin, Xin Yu

Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook that emotion is one of the key factors of authentic co-speech gesture generation. In this work, we propose EmotionGesture, a novel framework for synthesizing vivid and diverse emotional co-speech 3D gestures from audio. Considering emotion is often entangled with the rhythmic beat in speech audio, we first develop an Emotion-Beat Mining module (EBM) to extract the emotion and audio beat features as well as model their correlation via a transcript-based visual-rhythm alignment. Then, we propose an initial pose based Spatial-Temporal Prompter (STP) to generate future gestures from the given initial poses. STP effectively models the spatial-temporal correlations between the initial poses and the future gestures, thus producing the spatial-temporal coherent pose prompt. Once we obtain pose prompts, emotion, and audio beat features, we will generate 3D co-speech gestures through a transformer architecture. However, considering the poses of existing datasets often contain jittering effects, this would lead to generating unstable gestures. To address this issue, we propose an effective objective function, dubbed Motion-Smooth Loss. Specifically, we model motion offset to compensate for jittering ground-truth by forcing gestures to be smooth. Last, we present an emotion-conditioned VAE to sample emotion features, enabling us to generate diverse emotional results. Extensive experiments demonstrate that our framework outperforms the state-of-the-art, achieving vivid and diverse emotional co-speech 3D gestures. Our code and dataset will be released at the project page: https://xingqunqi-lab.github.io/Emotion-Gesture-Web/

📄 PDF Abstract BibTeX arXiv:2305.18891

Code (1)

xingqunqi-lab/emotiongestures 공식 구현 pytorch

Tasks

Gesture GenerationRhythm

Similar Papers 제목 키워드 기반

EmoFace: Audio-driven Emotional 3D Face Animation

2024-07-17 · Chang Liu, Qunfen Lin, Zijiao Zeng, Ye Pan

Audio-driven emotional 3D face animation aims to generate emotionally expressive talking heads with synchronized lip movements. However, previous research has often overlooked the influence of diverse emotions on facial …

3D Face Animation

Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head Animation

2025-01-01 · CVPR 2025 1 · Xiumei Xie, Zikai Huang, Wenhao Xu, Peng Xiao 외

Singing is a vital form of human emotional expression and social interaction, distinguished from speech by its richer emotional nuances and freer expressive style. Thus, investigating 3D facial animation driven by si…

Emotional Face-to-Speech

2025-02-03 · Jiaxin Ye, Boyuan Cao, Hongming Shan

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive lang…

Audio-Driven Emotional Video Portraits

2021-04-15 · CVPR 2021 1 · Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu 외

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important feat…

DisentanglementFace Generation

Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech

2025-10-29 · Pedro Corrêa, João Lima, Victor Moreno, Lucas Ueda 외 arxiv

Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide ran…

Speech Emotion Recognition