paper-with-me

Papers

Learning Predictive Visuomotor Coordination

2025-03-30 · Wenqi Jia, Bolin Lai, Miao Liu, Danfei Xu, James M. Rehg

Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive technologies. This work introduces a forecasting-based task for visuomotor modeling, where the goal is to predict head pose, gaze, and upper-body motion from egocentric visual and kinematic observations. We propose a \textit{Visuomotor Coordination Representation} (VCR) that learns structured temporal dependencies across these multimodal signals. We extend a diffusion-based motion modeling framework that integrates egocentric vision and kinematic sequences, enabling temporally coherent and accurate visuomotor predictions. Our approach is evaluated on the large-scale EgoExo4D dataset, demonstrating strong generalization across diverse real-world activities. Our results highlight the importance of multimodal integration in understanding visuomotor coordination, contributing to research in visuomotor learning and human behavior modeling.

📄 PDF Abstract BibTeX arXiv:2503.23300

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generalization Through Hand-Eye Coordination: An Action Space for Learning Spatially-Invariant Visuomotor Control

2021-02-28 · Chen Wang, Rui Wang, Ajay Mandlekar, Li Fei-Fei 외

Imitation Learning (IL) is an effective framework to learn visuomotor skills from offline demonstration data. However, IL methods often fail to generalize to new scene configurations not covered by training data. On the …

Imitation LearningZero-shot Generalization

FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination

2025-12-04 · Chengyang He, Ge Sun, Yue Bai, Junkai Lu 외 arxiv

We present FoundAtion-model-guided decoupled LoCO-maNipulation visuomotor policies (FALCON), a framework for loco-manipulation that combines modular diffusion policies with a vision-language foundation model as the coord…

CUBic: Coordinated Unified Bimanual Perception and Control Framework

2026-05-13 · Xingyu Wang, Pengxiang Ding, Jingkai Xu, Donglin Wang 외 arxiv

Recent advances in visuomotor policy learning have enabled robots to perform control directly from visual inputs. Yet, extending such end-to-end learning from single-arm to bimanual manipulation remains challenging due t…

Learning to Autonomously Reach Objects with NICO and Grow-When-Required Networks

2022-10-14 · Nima Rahrakhshan, Matthias Kerzel, Philipp Allgeuer, Nicolas Duczek 외

The act of reaching for an object is a fundamental yet complex skill for a robotic agent, requiring a high degree of visuomotor control and coordination. In consideration of dynamic environments, a robot capable of auton…

Object

FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model

2026-03-11 · Xiaoxu Xu, Hao Li, Jinhui Ye, Yilun Chen 외 arxiv

Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geometry, effectively anticipating the future …