paper-with-me

홈 › Papers

M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis

2025-05-13 · Zhizhuo Yin, Yuk Hang Tsui, Pan Hui

Generating full-body human gestures encompassing face, body, hands, and global movements from audio is a valuable yet challenging task in virtual avatar creation. Previous systems focused on tokenizing the human gestures framewisely and predicting the tokens of each frame from the input audio. However, one observation is that the number of frames required for a complete expressive human gesture, defined as granularity, varies among different human gesture patterns. Existing systems fail to model these gesture patterns due to the fixed granularity of their gesture tokens. To solve this problem, we propose a novel framework named Multi-Granular Gesture Generator (M3G) for audio-driven holistic gesture generation. In M3G, we propose a novel Multi-Granular VQ-VAE (MGVQ-VAE) to tokenize motion patterns and reconstruct motion sequences from different temporal granularities. Subsequently, we proposed a multi-granular token predictor that extracts multi-granular information from audio and predicts the corresponding motion tokens. Then M3G reconstructs the human gestures from the predicted tokens using the MGVQ-VAE. Both objective and subjective experiments demonstrate that our proposed M3G framework outperforms the state-of-the-art methods in terms of generating natural and expressive full-body human gestures.

📄 PDF Abstract BibTeX arXiv:2505.08293

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture GenerationMotion Synthesis

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…

Similar Papers 제목 키워드 기반

Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation

2022-03-24 · CVPR 2022 1 · Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu 외

Generating speech-consistent body and gesture movements is a long-standing problem in virtual avatar creation. Previous studies often synthesize pose movement in a holistic manner, where poses of all joints are generated…

Contrastive LearningGesture Generation

EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model

2025-04-11 · Renda Li, Xiaohua Qi, Qiang Ling, Jun Yu 외

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions an…

Gesture GenerationVideo Generation

Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation

2023-03-16 · CVPR 2023 1 · Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian 외

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from …

DiversityGesture Generation

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

2025-08-19 · Shaoshu Yang, Zhe Kong, Feng Gao, Meng Cheng 외 arxiv

Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant…

Video Generation

Audio-Driven Co-Speech Gesture Video Generation

2022-12-05 · Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 외

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…

Video Generation