paper-with-me

Papers

EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next

2026-03-12 · Ye Pan, Chi Kit Wong, Yuanhuiyi Lyu, Hanqian Li, Jiahao Huo, Jiacheng Chen, Lutao Jiang, Xu Zheng, Xuming Hu arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable video reasoning capabilities across diverse tasks. However, their ability to understand human intent at a fine-grained level in egocentric videos remains largely unexplored. Existing benchmarks focus primarily on episode-level intent reasoning, overlooking the finer granularity of step-level intent understanding. Yet applications such as intelligent assistants, robotic imitation learning, and augmented reality guidance require understanding not only what a person is doing at each step, but also why and what comes next, in order to provide timely and context-aware support. To this end, we introduce EgoIntent, a step-level intent understanding benchmark for egocentric videos. It comprises 3,014 steps spanning 15 diverse indoor and outdoor daily-life scenarios, and evaluates models on three complementary dimensions: local intent (What), global intent (Why), and next-step plan (Next). Crucially, each clip is truncated immediately before the key outcome of the queried step (e.g., contact or grasp) occurs and contains no frames from subsequent steps, preventing future-frame leakage and enabling a clean evaluation of anticipatory step understanding and next-step planning. We evaluate 15 MLLMs, including both state-of-the-art closed-source and open-source models. Even the best-performing model achieves an average score of only 33.31 across the three intent dimensions, underscoring that step-level intent understanding in egocentric videos remains a highly challenging problem that calls for further investigation.

📄 PDF Abstract BibTeX arXiv:2603.12147

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visual Intention Grounding for Egocentric Assistants

2025-04-18 · Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 외

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- …

ObjectVisual Grounding

MM-Ego: Towards Building Egocentric Multimodal LLMs

2024-10-09 · Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen 외

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric …

Video Understanding

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

2026-05-19 · Yang Dai, Dian Jiao, Tianwei Lin, Wenqiao Zhang arxiv

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, trac…

EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding

2024-06-13 · Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng 외

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-bo…

Action ClassificationAction LocalizationAction Understanding

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

2026-05-12 · Patrick Knab, Orgest Xhelili, Inis Buzi, Drago Andres Guggiana Nilo 외 arxiv

Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physical processes remains difficult, as model…

Visual Question AnsweringActivity RecognitionObject LocalizationScene Understanding