paper-with-me

홈 › Papers

Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents

2025-09-15 · Anna Deichler, Siyang Wang, Simon Alexanderson, Jonas Beskow arxiv

One of the main goals of robotics and intelligent agent research is to enable natural communication with humans in physically situated settings. While recent work has focused on verbal modes such as language and speech, non-verbal communication is crucial for flexible interaction. We present a framework for generating pointing gestures in embodied agents by combining imitation and reinforcement learning. Using a small motion capture dataset, our method learns a motor control policy that produces physically valid, naturalistic gestures with high referential accuracy. We evaluate the approach against supervised learning and retrieval baselines in both objective metrics and a virtual reality referential game with human users. Results show that our system achieves higher naturalness and accuracy than state-of-the-art supervised models, highlighting the promise of imitation-RL for communicative gesture generation and its potential application to robots.

📄 PDF Abstract BibTeX arXiv:2509.12507

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningGesture Generation

Similar Papers 제목 키워드 기반

Interpreting and Generating Gestures with Embodied Human Computer Interactions

2020-09-18 · ACM IVA Workshop GENEA 2020 10 · Anonymous

In this paper, we discuss the role that gesture plays for an embodied intelligent virtual agent (IVA) in the context of multimodal task-oriented dialogues with a human. We have developed a simulation platform, VoxWorl…

Gesture Generation

Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

2025-04-13 · Atharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen 외

3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pur…

3D visual groundingData AugmentationVisual Grounding

Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

2025-01-01 · CVPR 2025 1 · Atharv Mahesh Mane, Dulanga Weerakoon, Vigneshwaran Subbaraju, Sougata Sen 외

3-Dimensional Embodied Reference Understanding (3DERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored …

3D visual groundingData AugmentationVisual Grounding

Chinese Whispers: A Multimodal Dataset for Embodied Language Grounding

2020-05-01 · LREC 2020 5 · Dimosthenis Kontogiorgos, Elena Sibirtseva, Joakim Gustafson

In this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. Using the concept of {`}Chinese Whispers{'}, an old children{'}s game, we employ a novel metho…

Transformer Network for Semantically-Aware and Speech-Driven Upper-Face Generation

2021-10-09 · Mireille Fares, Catherine Pelachaud, Nicolas Obin

We propose a semantically-aware speech driven model to generate expressive and natural upper-facial and head motion for Embodied Conversational Agents (ECA). In this work, we aim to produce natural and continuous head mo…

Face Generation