paper-with-me

Papers

Modeling Turn-Taking with Semantically Informed Gestures

2025-10-22 · Varsha Suresh, M. Hamza Mughal, Christian Theobalt, Vera Demberg arxiv

In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide complementary cues for modeling these transitions. To study this, we introduce DnD Gesture++, an extension of the multi-party DnD Gesture corpus enriched with 2,663 semantic gesture annotations spanning iconic, metaphoric, deictic, and discourse types. Using this dataset, we model turn-taking prediction through a Mixture-of-Experts framework integrating text, audio, and gestures. Experiments show that incorporating semantically guided gestures yields consistent performance gains over baselines, demonstrating their complementary role in multimodal turn-taking.

📄 PDF Abstract BibTeX arXiv:2510.19350

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pragmatic Frames Evoked by Gestures: A FrameNet Brasil Approach to Multimodality in Turn Organization

2025-09-11 · Helen de Andrade Abreu, Tiago Timponi Torrent, Ely Edison da Silva Matos arxiv

This paper proposes a framework for modeling multimodal conversational turn organization via the proposition of correlations between language and interactive gestures, based on analysis as to how pragmatic frames are con…

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

2024-03-26 · CVPR 2024 1 · Muhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli 외

Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utter…

Gesture Generation

ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input

2025-10-20 · Hendric Voss, Stefan Kopp arxiv

Human communication combines speech with expressive nonverbal cues such as hand gestures that serve manifold communicative functions. Yet, current generative gesture generation approaches are restricted to simple, repeti…

Gesture Generation

Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis

2024-12-09 · CVPR 2025 1 · M. Hamza Mughal, Rishabh Dabral, Merel C. J. Scholman, Vera Demberg 외

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems t…

Gesture GenerationRAGRetrievalRetrieval-augmented Generation

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

2026-05-28 · Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman, M. Hamza Mughal 외 arxiv

Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically meaningful gestures whose communicative i…

Gesture GenerationText Retrieval