paper-with-me

홈 › Papers

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

2026-08-13 · Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-sourced videos that lack inter-clip spatiotemporal continuity, making it difficult to assess whether models can maintain consistent memory across days or weeks of real-world experience. We introduce EgoMonth, the first month-level egocentric video understanding benchmark. EgoMonth comprises over 300 hours of first-person daily-life recordings from 20 participants spanning 20 to 120 days, paired with 1,443 human-crafted multiple-choice question-answer pairs. We design a cognitively grounded 14-task evaluation framework organized into three hierarchical cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. Evaluation of state-of-the-art open-source and closed-source MLLMs reveals that even the best-performing model, Gemini 2.5 Pro, achieves only 71.8% macro-average accuracy, remaining 22.4 percentage points below the corrected human baseline of 94.2%. Several models perform near or below the 25% chance level on tasks such as Route Reasoning, Cross-view Spatial Reasoning, and Direction Judgement, while even the strongest closed-source model remains substantially below human performance. These results indicate that current MLLMs function as lossy summarizers rather than faithful memorizers, highlighting the need for architectures with genuine long-term spatiotemporal memory.

📄 PDF Abstract BibTeX arXiv:2608.13113

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Wanderlust: Online Continual Object Detection in the Real World

2021-08-25 · ICCV 2021 10 · Jianren Wang, Xin Wang, Yue Shang-Guan, Abhinav Gupta

Online continual learning from data streams in dynamic environments is a critical direction in the computer vision field. However, realistic benchmarks and fundamental studies in this line are still missing. To bridge th…

Continual LearningObjectobject-detectionObject Detection

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

2025-03-28 · YuXuan Li, Vijay Veerabadran, Michael L. Iuzzolino, Brett D. Roads 외

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D datase…

BenchmarkingQuestion AnsweringVideo Question Answering

Delving Into Egocentric Actions

2015-06-01 · CVPR 2015 6 · Yin Li, Zhefan Ye, James M. Rehg

We address the challenging problem of recognizing the camera wearer's actions from videos captured by an egocentric camera. Egocentric videos encode a rich set of signals regarding the camera wearer, including head movem…

Action RecognitionTemporal Action Localization

AMEGO: Active Memory from long EGOcentric videos

2024-09-17 · Gabriele Goletto, Tushar Nagarajan, Giuseppe Averta, Dima Damen

Egocentric videos provide a unique perspective into individuals' daily experiences, yet their unstructured nature presents challenges for perception. In this paper, we introduce AMEGO, a novel approach aimed at enhancing…

Video Understanding

Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance

2025-12-15 · Francesco Ragusa, Michele Mazzamuto, Rosario Forte, Irene D'Ambra 외 arxiv

We present Ego-EXTRA, a video-language Egocentric Dataset for EXpert-TRAinee assistance. Ego-EXTRA features 50 hours of unscripted egocentric videos of subjects performing procedural activities (the trainees) while guide…