paper-with-me

Papers

G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding

2026-08-20 · Marko Haralović, Akash Ramakrishnan, Estefania Talavera Martinez arxiv

Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets. However, many first-person actions depend on a small number of hand-object interactions involving only a few relevant entities. We propose G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene. From sparsely sampled frames, G3Ego constructs action scene graphs from vision-language descriptions, grounded objects, and hand cues, and then prunes irrelevant entities using the camera wearer's gaze. The resulting graph embeddings are temporally aggregated for action recognition and anticipation. Unlike prior work that uses gaze primarily as an auxiliary modality or attention signal, G3Ego incorporates gaze directly into graph construction, producing efficient and interpretable representations focused on action-relevant interactions. Experiments on EGTEA Gaze+ and MECCANO show that G3Ego achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on computationally expensive video pretraining. These results demonstrate the effectiveness of gaze-guided graph representations for egocentric action understanding.

📄 PDF Abstract BibTeX arXiv:2608.20157

Code (0)

등록된 구현이 없습니다.

Tasks

Action UnderstandingAction Recognition

Similar Papers 제목 키워드 기반

In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting

2025-09-09 · Taiying Peng, Jiacheng Hua, Miao Liu, Feng Lu arxiv

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing us…

Video Question AnsweringGaze Estimation

Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding

2025-10-24 · Anupam Pani, Yanchao Yang arxiv

Eye gaze offers valuable cues about attention, short-term intent, and future actions, making it a powerful signal for modeling egocentric behavior. In this work, we propose a gaze-regularized framework that enhances VLMs…

Mutual Context Network for Jointly Estimating Egocentric Gaze and Actions

2019-01-07 · Yifei Huang, Zhenqiang Li, Minjie Cai, Yoichi Sato

In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context. Our assumption is that in the procedure of performing a manipulation task, what…

Action RecognitionGaze PredictionPredictionTemporal Action Localization

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

2025-12-01 · Daeun Lee, Subhojyoti Mukherjee, Branislav Kveton, Ryan A. Rossi 외 arxiv

Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realistic applications such as Augmented Reality (AR) glasses. While prior streaming…

Video Question Answering

Eyes on Target: Gaze-Aware Object Detection in Egocentric Video

2025-11-03 · Vishakha Lall, Yisi Liu arxiv

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework desig…

Object Detection