paper-with-me

Papers

Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis

2024-12-09 · CVPR 2025 1 · M. Hamza Mughal, Rishabh Dabral, Merel C. J. Scholman, Vera Demberg, Christian Theobalt

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically meaningful gestures. Therefore, we present RAG-Gesture, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures. Our neuro-explicit gesture generation approach is designed to produce semantic gestures grounded in interpretable linguistic knowledge. We achieve this by using explicit domain knowledge to retrieve exemplar motions from a database of co-speech gestures. Once retrieved, we then inject these semantic exemplar gestures into our diffusion-based gesture generation pipeline using DDIM inversion and retrieval guidance at the inference time without any need of training. Further, we propose a control paradigm for guidance, that allows the users to modulate the amount of influence each retrieval insertion has over the generated sequence. Our comparative evaluations demonstrate the validity of our approach against recent gesture generation approaches. The reader is urged to explore the results on our project page.

📄 PDF Abstract BibTeX arXiv:2412.06786

Code (1)

andypinxinliu/GestureLSM pytorch

Tasks

Gesture GenerationRAGRetrievalRetrieval-augmented Generation

Similar Papers 제목 키워드 기반

Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings

2022-10-04 · Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen 외

Automatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation. Previous systems mainly focus on generating gestures in an end-to-end manner, w…

Gesture GenerationRhythm

SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain

2025-03-26 · Nan Gao, Yihua Bao, Dongdong Weng, Jiayi Zhao 외

Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGe…

Gesture Generation

BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer

2023-09-07 · Kunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost 외

Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is diff…

DeepGesture: A conversational gesture synthesis system based on emotions and semantics

2025-07-03 · Thanh Hoang-Minh

Along with the explosion of large language models, improvements in speech synthesis, advancements in hardware, and the evolution of computer graphics, the current bottleneck in creating digital humans lies in generating …

Gesture GenerationMotion SynthesisSpeech SynthesisUnity

BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis

2022-03-10 · Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng 외

Achieving realistic, vivid, and human-like synthesized conversational gestures conditioned on multi-modal data is still an unsolved problem due to the lack of available datasets, models and standard evaluation metrics. T…

Gesture GenerationGesture Recognition