paper-with-me

홈 › Papers

EMoG: Synthesizing Emotive Co-speech 3D Gesture with Diffusion Model

2023-06-20 · Lianying Yin, Yijun Wang, Tianyu He, Jinming Liu, Wei Zhao, Bohan Li, Xin Jin, Jianxin Lin

Although previous co-speech gesture generation methods are able to synthesize motions in line with speech content, it is still not enough to handle diverse and complicated motion distribution. The key challenges are: 1) the one-to-many nature between the speech content and gestures; 2) the correlation modeling between the body joints. In this paper, we present a novel framework (EMoG) to tackle the above challenges with denoising diffusion models: 1) To alleviate the one-to-many problem, we incorporate emotion clues to guide the generation process, making the generation much easier; 2) To model joint correlation, we propose to decompose the difficult gesture generation into two sub-problems: joint correlation modeling and temporal dynamics modeling. Then, the two sub-problems are explicitly tackled with our proposed Joint Correlation-aware transFormer (JCFormer). Through extensive evaluations, we demonstrate that our proposed method surpasses previous state-of-the-art approaches, offering substantial superiority in gesture synthesis.

📄 PDF Abstract BibTeX arXiv:2306.11496

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingGesture Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

AMUSE: Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2024-06-01 · IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) 2024 6 · Chhatre K., Danecek R., Athanasiou N., Becherini G. 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm

Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion

2023-12-07 · CVPR 2024 1 · Kiran Chhatre, Radek Daněček, Nikos Athanasiou, Giorgio Becherini 외

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output anim…

Gesture GenerationRhythm

EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model

2025-04-11 · Renda Li, Xiaohua Qi, Qiang Ling, Jun Yu 외

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions an…

Gesture GenerationVideo Generation

GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents

2023-03-26 · Tenglong Ao, Zeyi Zhang, Libin Liu

The automatic generation of stylized co-speech gestures has recently received increasing attention. Previous systems typically allow style control via predefined text labels or example motion clips, which are often not f…

Contrastive LearningGesture Generationmodel

Real-time Gesture Animation Generation from Speech for Virtual Human Interaction

2022-08-05 · Manuel Rebol, Christian Gütl, Krzysztof Pietroszek

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amo…