paper-with-me

Papers

SemConFlow: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching

2026-03-27 · Lanmiao Liu, Esam Ghaleb, Aslı Özyürek, Zerrin Yumak arxiv

While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challenge. Existing approaches rely on external semantic retrieval methods, which limit their generalisation capability due to dependency on predefined linguistic rules. Flow-matching-based methods produce promising results; however, the network is optimised using only semantically congruent samples without exposure to negative examples, leading to learning rhythmic gestures rather than sparse motion, such as iconic and metaphoric gestures. Furthermore, by modelling body parts in isolation, the majority of methods fail to maintain crossmodal consistency. We introduce a Contrastive Flow Matching-based co-speech gesture generation model that uses mismatched audio-text conditions as negatives, training the velocity field to follow the correct motion trajectory while repelling semantically incongruent trajectories. Our model ensures cross-modal coherence by embedding text, audio, and holistic motion into a composite latent space via cosine and contrastive objectives. Extensive experiments and a user study demonstrate that our proposed approach outperforms state-of-the-art methods on two datasets, BEAT2 and SHOW.

📄 PDF Abstract BibTeX arXiv:2603.26553

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic RetrievalGesture Generation

Similar Papers 제목 키워드 기반

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

2026-05-25 · Ferdinand Paar, Lanmiao Liu, Aslı Özyürek, Serge Thill 외 arxiv

Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lexically grounded semantic gestures with frequent prosody-aligned beat…

Gesture Generation

3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control

2026-01-26 · Xuanmeng Sha, Liyun Zhang, Tomohiro Mashita, Naoya Chiba 외 arxiv

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing…

Gesture Generation

SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning

2025-07-25 · Lanmiao Liu, Esam Ghaleb, Aslı Özyürek, Zerrin Yumak arxiv

Creating a virtual avatar with semantically coherent gestures that are aligned with speech is a challenging task. Existing gesture generation research mainly focused on generating rhythmic beat gestures, neglecting the s…

Gesture Generation

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

2024-12-21 · Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang 외

A good co-speech motion generation cannot be achieved without a careful integration of common rhythmic motion and rare yet essential semantic motion. In this work, we propose SemTalk for holistic co-speech motion generat…

Gesture GenerationMotion GenerationRhythm

Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation

2022-03-24 · CVPR 2022 1 · Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu 외

Generating speech-consistent body and gesture movements is a long-standing problem in virtual avatar creation. Previous studies often synthesize pose movement in a holistic manner, where poses of all joints are generated…

Contrastive LearningGesture Generation