paper-with-me

Papers

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

2025-07-21 · Ian Chuang, Jinyu Zou, Andrew Lee, Dechen Gao, Iman Soltani arxiv

Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform processing of raw camera images. In this work, we explore how incorporating human-like active gaze into robotic policies can enhance efficiency and robustness. We develop GIAVA (Gaze Integrated Active-Vision ALOHA), a robot vision system that emulates human head and neck movement, and gaze adjustment for foveated processing. Extending the AV-ALOHA robot platform, we introduce a framework for simultaneously collecting eye-tracking, perspective control, and robot manipulation demonstration data from a human operator. We also open-source a simulation benchmark and dataset for training robot policies that incorporate human gaze. Inspired by recent work in foveated image segmentation and given the widespread use of Vision Transformers (ViTs) in robot learning, we integrate gaze information into ViTs using a foveated patch tokenization scheme. Compared to uniform patch tokenization, this significantly reduces the number of tokens, and thus computation. Our results show that our method for foveated robot vision drastically reduces computational overhead, and enhances robustness to background distractors. Notably, on certain high-precision tasks, foveated vision also improves performance, as reflected in higher success rates. Together, these findings suggest that human-inspired foveated visual processing offers untapped potential and should be further considered as a useful inductive bias in robotic vision systems. https://soltanilara.github.io/giava/

📄 PDF Abstract BibTeX arXiv:2507.15833

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationImage Segmentation

Similar Papers 제목 키워드 기반

Impact of Gaze-Based Interaction and Augmentation on Human-Robot Collaboration in Critical Tasks

2025-08-10 · Ayesha Jena, Stefan Reitmann, Elin Anna Topp arxiv

We present a user study analyzing head-gaze-based robot control and foveated visual augmentation in a simulated search-and-rescue task. Results show that foveated augmentation significantly improves task performance, red…

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

2026-08-17 · Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci 외 arxiv

Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do be…

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

2026-05-18 · Shravan Murlidaran, Ziqi Wen, Sana Shehabi, Miguel P. Eckstein arxiv

When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on people, text, objects being gazed at or grasped, and semantically meani…

Scene Understanding

A3FR: Agile 3D Gaussian Splatting with Incremental Gaze Tracked Foveated Rendering in Virtual Reality

2025-07-05 · Shuo Xin, Haiyu Wang, Sai Qian Zhang arxiv

Virtual reality (VR) significantly transforms immersive digital interfaces, greatly enhancing education, professional practices, and entertainment by increasing user engagement and opening up new possibilities in various…

FovealNet: Advancing AI-Driven Gaze Tracking Solutions for Optimized Foveated Rendering System Performance in Virtual Reality

2024-12-12 · Wenxuan Liu, Monde Duinkharjav, Qi Sun, Sai Qian Zhang

Leveraging real-time eye-tracking, foveated rendering optimizes hardware efficiency and enhances visual quality virtual reality (VR). This approach leverages eye-tracking techniques to determine where the user is looking…