paper-with-me

홈 › Papers

Co-speech Gesture Video Generation via Motion-Based Graph Retrieval

2025-12-02 · Yafei Song, Peng Zhang, Bang Zhang arxiv

Synthesizing synchronized and natural co-speech gesture videos remains a formidable challenge. Recent approaches have leveraged motion graphs to harness the potential of existing video data. To retrieve an appropriate trajectory from the graph, previous methods either utilize the distance between features extracted from the input audio and those associated with the motions in the graph or embed both the input audio and motion into a shared feature space. However, these techniques may not be optimal due to the many-to-many mapping nature between audio and gestures, which cannot be adequately addressed by one-to-one mapping. To alleviate this limitation, we propose a novel framework that initially employs a diffusion model to generate gesture motions. The diffusion model implicitly learns the joint distribution of audio and motion, enabling the generation of contextually appropriate gestures from input audio sequences. Furthermore, our method extracts both low-level and high-level features from the input audio to enrich the training process of the diffusion model. Subsequently, a meticulously designed motion-based retrieval algorithm is applied to identify the most suitable path within the graph by assessing both global and local similarities in motion. Given that not all nodes in the retrieved path are sequentially continuous, the final step involves seamlessly stitching together these segments to produce a coherent video output. Experimental results substantiate the efficacy of our proposed method, demonstrating a significant improvement over prior approaches in terms of synchronization accuracy and naturalness of generated gestures.

📄 PDF Abstract BibTeX arXiv:2512.02576

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Audio-Driven Co-Speech Gesture Video Generation

2022-12-05 · Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 외

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…

Video Generation

MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation

2025-05-29 · Siyuan Wang, Jiawei Liu, Wei Wang, Yeying Jin 외

Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of different parts of the body in terms of amplitude of motion, audio rele…

Motion GenerationVideo Generation

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

2024-04-02 · CVPR 2024 1 · Xu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin 외

Co-speech gestures, if presented in the lively form of videos, can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons, resulting in the omission …

Video Generation

Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation

2025-02-11 · Pinxin Liu, Pengfei Zhang, Hyeongwoo Kim, Pablo Garrido 외

Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accuratel…

Gesture GenerationVideo Generation

Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement

2025-01-01 · CVPR 2025 1 · Xinjie Li, Ziyi Chen, Xinlu Yu, Iek-Heng Chu 외

Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, …

Gesture GenerationMotion GenerationVideo Generation