paper-with-me

Papers

Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models

2025-07-27 · Bohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding, Kun Zhou arxiv

The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels derived from motion examples, these approaches often compromise the rich details present in the original motion examples. We present MECo, a framework for motion-example-controlled co-speech gesture generation by leveraging large language models (LLMs). Our method capitalizes on LLMs' comprehension capabilities through fine-tuning to simultaneously interpret speech audio and motion examples, enabling the synthesis of gestures that preserve example-specific characteristics while maintaining speech congruence. Departing from conventional pseudo-labeling paradigms, we position motion examples as explicit query contexts within the prompt structure to guide gesture generation. Experimental results demonstrate state-of-the-art performance across three metrics: Fréchet Gesture Distance (FGD), motion diversity, and example-gesture similarity. Furthermore, our framework enables granular control of individual body parts and accommodates diverse input modalities including motion clips, static poses, human video sequences, and textual descriptions. Our code, pre-trained models, and videos are available at https://robinwitch.github.io/MECo-Page.

📄 PDF Abstract BibTeX arXiv:2507.20220

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture Generation

Similar Papers 제목 키워드 기반

ZeroEGGS: Zero-shot Example-based Gesture Generation from Speech

2022-09-15 · Saeed Ghorbani, Ylva Ferstl, Daniel Holden, Nikolaus F. Troje 외

We present ZeroEGGS, a neural network framework for speech-driven gesture generation with zero-shot style control by example. This means style can be controlled via only a short example motion clip, even for motion style…

Gesture Generation

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

2025-11-28 · Fengyi Fang, Sicheng Yang, Wenming Yang arxiv

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Exist…

Gesture Generation

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

2026-05-25 · Ferdinand Paar, Lanmiao Liu, Aslı Özyürek, Serge Thill 외 arxiv

Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lexically grounded semantic gestures with frequent prosody-aligned beat…

Gesture Generation

Audio-Driven Co-Speech Gesture Video Generation

2022-12-05 · Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 외

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…

Video Generation

Moving fast and slow: Analysis of representations and post-processing in speech-driven automatic gesture generation

2020-07-16 · Taras Kucherenko, Dai Hasegawa, Naoshi Kaneko, Gustav Eje Henter 외

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for …

Gesture GenerationRepresentation Learning