paper-with-me

Papers

Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement

2025-01-01 · CVPR 2025 1 · Xinjie Li, Ziyi Chen, Xinlu Yu, Iek-Heng Chu, Peng Chang, Jing Xiao

Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS keypoint movements, which often lead to artifacts like blurry hands and distorted fingers. In response to these challenges, we present the Implicit Motion-Audio Entanglement (IMAE) method for co-speech gesture video generation. IMAE strengthens audio control by entangling implicit motion parameters, including pose and expression, with audio inputs. Our method utilizes a two-branch framework that combines an audio-to-motion generation branch with a video diffusion branch, enabling realistic gesture generation without requiring additional inputs during inference. To improve training efficiency, we propose a two-stage slow-fast training strategy that balances memory constraints while facilitating the learning of meaningful gestures from long frame sequences.Extensive experimental results demonstrate that our method achieves state-of-the-art performance across multiple metrics.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture GenerationMotion GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Audio-Driven Co-Speech Gesture Video Generation

2022-12-05 · Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 외

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…

Video Generation

Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models

2025-07-27 · Bohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding 외 arxiv

The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels …

Gesture Generation

MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation

2025-05-29 · Siyuan Wang, Jiawei Liu, Wei Wang, Yeying Jin 외

Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of different parts of the body in terms of amplitude of motion, audio rele…

Motion GenerationVideo Generation

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

2024-04-02 · CVPR 2024 1 · Xu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin 외

Co-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons, resulting in the omission …

Video Generation

Co-speech Gesture Video Generation via Motion-Based Graph Retrieval

2025-12-02 · Yafei Song, Peng Zhang, Bang Zhang arxiv

Synthesizing synchronized and natural co-speech gesture videos remains a formidable challenge. Recent approaches have leveraged motion graphs to harness the potential of existing video data. To retrieve an appropriate tr…

Video Generation