paper-with-me

Papers

This&That: Language-Gesture Controlled Video Generation for Robot Planning

2024-07-08 · Boyang Wang, Nikhil Sridhar, Chao Feng, Mark Van der Merwe, Adam Fishman, Nima Fazeli, Jeong Joon Park

Clear, interpretable instructions are invaluable when attempting any complex task. Good instructions help to clarify the task and even anticipate the steps needed to solve it. In this work, we propose a robot learning framework for communicating, planning, and executing a wide range of tasks, dubbed This&That. This&That solves general tasks by leveraging video generative models, which, through training on internet-scale data, contain rich physical and semantic context. In this work, we tackle three fundamental challenges in video-based planning: 1) unambiguous task communication with simple human instructions, 2) controllable video generation that respects user intent, and 3) translating visual plans into robot actions. This&That uses language-gesture conditioning to generate video predictions, as a succinct and unambiguous alternative to existing language-only methods, especially in complex and uncertain environments. These video predictions are then fed into a behavior cloning architecture dubbed Diffusion Video to Action (DiVA), which outperforms prior state-of-the-art behavior cloning and video-based planning methods by substantial margins.

📄 PDF Abstract BibTeX arXiv:2407.05530

Code (0)

등록된 구현이 없습니다.

Tasks

Task PlanningVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models

2025-07-27 · Bohong Chen, Yumeng Li, Youyi Zheng, Yao-Xiang Ding 외 arxiv

The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels …

Gesture Generation

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

2026-05-21 · Wenxuan Guo, Ziyuan Li, Meng Zhang, Yichen Liu 외 arxiv

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily rely on textual instructions and struggle …

Robot Manipulation

Teaching Arm and Head Gestures to a Humanoid Robot through Interactive Demonstration and Spoken Instruction

2021-06-01 · ACL (mmsr, IWCS) 2021 6 · Michael Brady, Han Du

We describe work in progress for training a humanoid robot to produce iconic arm and head gestures as part of task-oriented dialogue interaction. This involves the development and use of a multimodal dialog manager for n…

Gesture Recognition

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

2025-11-28 · Fengyi Fang, Sicheng Yang, Wenming Yang arxiv

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Exist…

Gesture Generation

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

2024-03-26 · CVPR 2024 1 · Muhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli 외

Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utter…

Gesture Generation