paper-with-me

홈 › Papers

Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation

2026-03-24 · Anupam Pani, Yanchao Yang arxiv

Despite advances in Vision-Language-Action (VLA) models, robotic manipulation struggles with fine-grained tasks because current models lack mechanisms for active visual attention allocation. Human gaze naturally encodes intent, planning, and execution patterns -- offering a powerful supervisory signal for guiding robot perception. We introduce a gaze-regularized training framework that aligns VLA models' internal attention with human visual patterns without architectural modifications or inference-time overhead. Our method transforms temporally aggregated gaze heatmaps into patch-level distributions and regularizes the transformer's attention through KL divergence, creating an inductive bias toward task-relevant features while preserving deployment efficiency. When integrated into existing VLA architectures, our approach yields 4-12% improvements across manipulation benchmarks. The gaze-regularized models reach equivalent performance with fewer training steps and maintain robustness under lighting variations and sensor noise. Beyond performance metrics, the learned attention patterns produce interpretable visualizations that mirror human strategies, enhancing trust in robotic systems. Moreover, our framework requires no eye-tracking equipment and applies directly to existing datasets. These results demonstrate that human perceptual priors can significantly accelerate robot learning while improving both task performance and system interpretability.

📄 PDF Abstract BibTeX arXiv:2603.23202

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gaze-Regularized VLMs for Ego-Centric Behavior Understanding

2026-03-24 · Anupam Pani, Yanchao Yang arxiv

Eye gaze, encompassing fixations and saccades, provides critical insights into human intentions and future actions. This study introduces a gaze-regularized framework that enhances Vision Language Models (VLMs) for egoce…

Intent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models

2026-01-08 · Tracey Yee Hsin Tay, Xu Yan, Jonathan Ouyang, Daniel Wu 외 arxiv

Designing intuitive interfaces for robotic control remains a central challenge in enabling effective human-robot interaction, particularly in assistive care settings. Eye gaze offers a fast, non-intrusive, and intent-ric…

Gazebo Plants: Simulating Plant-Robot Interaction with Cosserat Rods

2024-02-04 · Junchen Deng, Samhita Marri, Jonathan Klein, Wojtek Pałubicki 외

Robotic harvesting has the potential to positively impact agricultural productivity, reduce costs, improve food quality, enhance sustainability, and to address labor shortage. In the rapidly advancing field of agricultur…

Image Segmentationobject-detectionObject DetectionSemantic Segmentation

RaycastGrasp: Eye-Gaze Interaction with Wearable Devices for Robotic Manipulation

2025-10-25 · Zitiantao Lin, Yongpeng Sang, Yang Ye arxiv

Robotic manipulators are increasingly used to assist individuals with mobility impairments in object retrieval. However, the predominant joystick-based control interfaces can be challenging due to high precision requirem…

Intent RecognitionObject Recognition

ETH-XGaze: A Large Scale Dataset for Gaze Estimation under Extreme Head Pose and Gaze Variation

2020-07-31 · ECCV 2020 8 · Xucong Zhang, Seonwook Park, Thabo Beeler, Derek Bradley 외

Gaze estimation is a fundamental task in many applications of computer vision, human computer interaction and robotics. Many state-of-the-art methods are trained and tested on custom datasets, making comparison across me…

Gaze Estimation