paper-with-me

홈 › Papers

ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis

2025-03-09 · Xukun Zhou, Fengxin Li, Ming Chen, Yan Zhou, Pengfei Wan, Di Zhang, Hongyan Liu, Jun He, Zhaoxin Fan

Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce gestures that are coarse, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a novel retrieval-enhanced diffusion framework with three key designs: (1) a Motion Base Construction, which builds a gesture library using training dataset; (2) a Motion Retrieval Module, employing constrative learning and momentum distillation for fine-grained reference poses retreiving; and (3) a Precision Control Module, integrating partial masking and stochastic masking to enable flexible and fine-grained control. Experimental evaluations on BEAT2 demonstrate that ExGes reduces Fr\'echet Gesture Distance by 6.2\% and improves motion diversity by 5.3\% over EMAGE, with user studies revealing a 71.3\% preference for its naturalness and semantic relevance. Code will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2503.06499

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityRetrieval

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Library 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

2025-09-20 · Sirui Wang, Andong Chen, Tiejun Zhao arxiv

Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predefined labels, reference audio, or natural…

Speech Synthesis

Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars

2026-04-16 · Yicheng Gong, Jiawei Zhang, Liqiang Liu, Yanwen Wang 외 arxiv

We present a framework for explicit emotion control in feed-forward, single-image 3D head avatar reconstruction. Unlike existing pipelines where emotion is implicitly entangled with geometry or appearance, we treat emoti…

3DXTalker: Unifying Identity, Lip Sync, Emotion, and Spatial Dynamics in Expressive 3D Talking Avatars

2026-02-11 · Zhongju Wang, Zhenhong Sun, Beier Wang, Yifu Wang 외 arxiv

Audio-driven 3D talking avatar generation is increasingly important in virtual communication, digital humans, and interactive media, where avatars must preserve identity, synchronize lip motion with speech, express emoti…

Semantic Co-Speech Gesture Synthesis and Real-Time Control for Humanoid Robots

2025-12-19 · Gang Zhang arxiv

We present an innovative end-to-end framework for synthesizing semantically meaningful co-speech gestures and deploying them in real-time on a humanoid robot. This system addresses the challenge of creating natural, expr…

Gesture Generation

Bidirectionally Deformable Motion Modulation For Video-based Human Pose Transfer

2023-07-15 · ICCV 2023 1 · Wing-Yin Yu, Lai-Man Po, Ray C. C. Cheung, Yuzhi Zhao 외

Video-based human pose transfer is a video-to-video generation task that animates a plain source human image based on a series of target human poses. Considering the difficulties in transferring highly structural pattern…

motion predictionPose TransferStyle TransferVideo Generation