paper-with-me

홈 › Papers

Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning

2021-07-31 · Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh Manocha

We present a generative adversarial network to synthesize 3D pose sequences of co-speech upper-body gestures with appropriate affective expressions. Our network consists of two components: a generator to synthesize gestures from a joint embedding space of features encoded from the input speech and the seed poses, and a discriminator to distinguish between the synthesized pose sequences and real 3D pose sequences. We leverage the Mel-frequency cepstral coefficients and the text transcript computed from the input speech in separate encoders in our generator to learn the desired sentiments and the associated affective cues. We design an affective encoder using multi-scale spatial-temporal graph convolutions to transform 3D pose sequences into latent, pose-based affective features. We use our affective encoder in both our generator, where it learns affective features from the seed poses to guide the gesture synthesis, and our discriminator, where it enforces the synthesized gestures to contain the appropriate affective expressions. We perform extensive evaluations on two benchmark datasets for gesture synthesis from the speech, the TED Gesture Dataset and the GENEA Challenge 2020 Dataset. Compared to the best baselines, we improve the mean absolute joint error by 10--33%, the mean acceleration difference by 8--58%, and the Fr\'echet Gesture Distance by 21--34%. We also conduct a user study and observe that compared to the best current baselines, around 15.28% of participants indicated our synthesized gestures appear more plausible, and around 16.32% of participants felt the gestures had more appropriate affective expressions aligned with the speech.

📄 PDF Abstract BibTeX arXiv:2108.00262

Code (1)

UttaranB127/speech2affective_gestures 공식 구현 pytorch

Tasks

Generative Adversarial NetworkGesture Generation

Similar Papers 제목 키워드 기반

Real-time Gesture Animation Generation from Speech for Virtual Human Interaction

2022-08-05 · Manuel Rebol, Christian Gütl, Krzysztof Pietroszek

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amo…

AMUSE: Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2024-06-01 · IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) 2024 6 · Chhatre K., Danecek R., Athanasiou N., Becherini G. 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm

Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2023-12-07 · CVPR 2024 1 · Kiran Chhatre, Radek Daněček, Nikos Athanasiou, Giorgio Becherini 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm

Co-Speech Gesture Synthesis by Reinforcement Learning With Contrastive Pre-Trained Rewards

2023-01-01 · CVPR 2023 1 · Mingyang Sun, Mengchen Zhao, Yaqing Hou, Minglei Li 외

There is a growing demand of automatically synthesizing co-speech gestures for virtual characters. However, it remains a challenge due to the complex relationship between input speeches and target gestures. Most exis…

reinforcement-learningReinforcement Learning (RL)

EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation

2023-05-30 · Xingqun Qi, Chen Liu, Lincheng Li, Jie Hou 외

Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook that emoti…

Gesture GenerationRhythm