paper-with-me

홈 › Papers

Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos

2026-04-04 · Daniele Materia, Francesco Ragusa, Giovanni Maria Farinella arxiv

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such capabilities requires to approach several complex challenges. This work addresses the problem of human-object interaction anticipation in Egocentric Vision using Vision Large Language Models (VLLMs). We tackle key limitations in existing approaches by improving visual grounding capabilities through Set-of-Mark prompting and understanding user intent via the trajectory formed by the user's most recent gaze fixations. To effectively capture the temporal dynamics immediately preceding the interaction, we further introduce a novel inverse exponential sampling strategy for input video frames. Experiments conducted on the egocentric dataset HD-EPIC demonstrate that our method surpasses state-of-the-art approaches for the considered task, showing its model-agnostic nature.

📄 PDF Abstract BibTeX arXiv:2604.03667

Code (0)

등록된 구현이 없습니다.

Tasks

Human-Object Interaction AnticipationVisual Grounding

Similar Papers 제목 키워드 기반

Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention

2023-03-27 · CVPR 2023 1 · Sounak Mondal, Zhibo Yang, Seoyoung Ahn, Dimitris Samaras 외

Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predi…

DecoderGaze PredictionLanguage ModellingPrediction+2

Naming, Describing, and Quantifying Visual Objects in Humans and LLMs

2024-03-11 · Alberto Testoni, Juell Sprott, Sandro Pezzelle

While human speakers use a variety of different expressions when describing the same object in an image, giving rise to a distribution of plausible labels driven by pragmatic constraints, the extent to which current Visi…

Hierarchical structure understanding in complex tables with VLLMs: a benchmark and experiments

2025-11-11 · Luca Bindini, Simone Giovannini, Simone Marinai, Valeria Nardoni 외 arxiv

This work investigates the ability of Vision Large Language Models (VLLMs) to understand and interpret the structure of tables in scientific articles. Specifically, we explore whether VLLMs can infer the hierarchical str…

Prompt Engineering

Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities

2024-06-12 · CVPR 2025 1 · Michele Mazzamuto, Antonino Furnari, Yoichi Sato, Giovanni Maria Farinella

We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze signals. While traditional methods rely on manually labeled mistakes, our approach d…

Gaze PredictionMistake Detection

Enhancing Reusability of Learned Skills for Robot Manipulation via Gaze and Bottleneck

2025-02-25 · Ryo Takizawa, Izumi Karino, Koki Nakagawa, Yoshiyuki Ohmura 외

Autonomous agents capable of diverse object manipulations should be able to acquire a wide range of manipulation skills with high reusability. Although advances in deep learning have made it increasingly feasible to repl…

Imitation LearningObjectRobot Manipulation