paper-with-me

Papers

Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

2026-05-28 · Kuangji Zuo, Gen Li, Bofan Lyu, Yanshuo Lu, Boyu Ma, Shijia Han, Xinyu Zhou, Xichen Yuan, Chuhao Zhou, Jiaqi Bai, Geng Li, Jianfei Yang arxiv

Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to describe which exact object to interact with among similar candidates, where to act on the object, or how the target may change during execution. To address this limitation, we propose Gaze2Act, a novel VLA framework that leverages human gaze as a dynamic and intuitive intent signal for complex interactive manipulation. Gaze2Act first bridges the ego-exo view gap by mapping first-person gaze into the robot's perspective through cross-view semantic matching, producing both an object mask and a gaze point for coarse-to-fine target specification. These cues are then integrated into the policy through perception-level prompting and action-level conditioning, allowing the robot to attend to relevant regions and execute precise interactions under dynamic intent. In a systematic evaluation across seven task categories and 16 real-robot tasks on a Unitree G1 humanoid, Gaze2Act achieves state-of-the-art performance in both intent accuracy and task success rate. It notably outperforms baselines in object disambiguation, fine-grained interaction, and dynamic intent steering. These results demonstrate that human gaze provides a natural, low-burden, and highly expressive modality for human-in-the-loop VLA control.

📄 PDF Abstract BibTeX arXiv:2605.30282

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation

2026-02-04 · Jia Li, Wenjie Zhao, Shijian Deng, Bolin Lai 외 arxiv

Online egocentric gaze estimation predicts where a camera wearer is looking from first-person video using only past and current frames, a task essential for augmented reality and assistive technologies. Unlike third-pers…

Gaze Estimation

Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes

2026-03-30 · Luke Palmer, Petar Palasek, Hazem Abdelkawy arxiv

Accurately modelling human attention is essential for numerous computer vision applications, particularly in the domain of automotive safety. Existing methods typically collapse gaze into saliency maps or scanpaths, trea…

VL4Gaze: Unleashing Vision-Language Models for Gaze Following

2025-12-23 · Shijing Wang, Chaoqun Cui, Yaping Huang, Hyung Jin Chang 외 arxiv

Human gaze provides essential cues for interpreting attention, intention, and social interaction in visual scenes, yet gaze understanding remains largely unexplored in current vision-language models (VLMs). While recent …

Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following

2026-05-21 · Shijing Wang, Yaping Huang, Chaoqun Cui, David Wong 외 arxiv

Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation models (VFMs) have demonstrated strong performance on this task, enabling…

Scene Understanding

Humanizing Robot Gaze Shifts: A Framework for Natural Gaze Shifts in Humanoid Robots

2026-02-25 · Jingchao Wei, Jingkai Qin, Yuxiao Cao, Jingcheng Huang 외 arxiv

Leveraging auditory and visual feedback for attention reorientation is essential for natural gaze shifts in social interaction. However, enabling humanoid robots to perform natural and context-appropriate gaze shifts in …