paper-with-me

홈 › Papers

EgoTaskQA: Understanding Human Tasks in Egocentric Videos

2022-10-08 · Baoxiong Jia, Ting Lei, Song-Chun Zhu, Siyuan Huang

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on object states (i.e., state changes), and their causal dependencies. These challenges are further aggravated by the natural parallelism from multi-tasking and partial observations in multi-agent collaboration. Most prior works leverage action localization or future prediction as an indirect metric for evaluating such task understanding from videos. To make a direct evaluation, we introduce the EgoTaskQA benchmark that provides a single home for the crucial dimensions of task understanding through question-answering on real-world egocentric videos. We meticulously design questions that target the understanding of (1) action dependencies and effects, (2) intents and goals, and (3) agents' beliefs about others. These questions are divided into four types, including descriptive (what status?), predictive (what will?), explanatory (what caused?), and counterfactual (what if?) to provide diagnostic analyses on spatial, temporal, and causal understandings of goal-oriented tasks. We evaluate state-of-the-art video reasoning models on our benchmark and show their significant gaps between humans in understanding complex goal-oriented egocentric videos. We hope this effort will drive the vision community to move onward with goal-oriented video understanding and reasoning.

📄 PDF Abstract BibTeX arXiv:2210.03929

Code (1)

Buzz-Beater/EgoTaskQA 공식 구현 pytorch

Tasks

Action LocalizationcounterfactualDescriptiveDiagnosticFuture predictionQuestion AnsweringVideo Understanding

Similar Papers 제목 키워드 기반

DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering

2025-10-23 · Jiayi Zou, Chaofan Chen, Bing-Kun Bao, Changsheng Xu arxiv

Egocentric Video Question Answering (Egocentric VideoQA) plays an important role in egocentric video understanding, which refers to answering questions based on first-person videos. Although existing methods have made pr…

Video Question Answering

MECCANO: A Multimodal Egocentric Dataset for Humans Behavior Understanding in the Industrial-like Domain

2022-09-19 · Francesco Ragusa, Antonino Furnari, Giovanni Maria Farinella

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person…

Action AnticipationAction RecognitionHuman-Object Interaction Detection

Self-Supervised Object Detection from Egocentric Videos

2023-01-01 · ICCV 2023 1 · Peri Akiva, Jing Huang, Kevin J Liang, Rama Kovvuri 외

Understanding the visual world from the perspective of humans (egocentric) has been a long-standing challenge in computer vision. Egocentric videos exhibit high scene complexity and irregular motion flows compared to…

Class-agnostic Object DetectionObjectobject-detectionObject Detection+2

The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain

2020-10-12 · Francesco Ragusa, Antonino Furnari, Salvatore Livatino, Giovanni Maria Farinella

Wearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in ego…

Action RecognitionActive Object DetectionHuman-Object Interaction DetectionObject+3

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

2025-08-18 · Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to…