paper-with-me

홈 › Papers

EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

2023-08-17 · NeurIPS 2023 11 · Karttikeya Mangalam, Raiymbek Akshulakov, Jitendra Malik

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human curated multiple choice question answer pairs, spanning over 250 hours of real video data, covering a very broad range of natural human activity and behavior. For each question, EgoSchema requires the correct answer to be selected between five given options based on a three-minute-long video clip. While some prior works have proposed video datasets with long clip lengths, we posit that merely the length of the video clip does not truly capture the temporal difficulty of the video task that is being considered. To remedy this, we introduce temporal certificate sets, a general notion for capturing the intrinsic temporal understanding length associated with a broad range of video understanding tasks & datasets. Based on this metric, we find EgoSchema to have intrinsic temporal lengths over 5.7x longer than the second closest dataset and 10x to 100x longer than any other video understanding dataset. Further, our evaluation of several current state-of-the-art video and language models shows them to be severely lacking in long-term video understanding capabilities. Even models with several billions of parameters achieve QA accuracy less than 33% (random is 20%) on the EgoSchema multi-choice question answering task, while humans achieve about 76% accuracy. We posit that \name{}{}, with its long intrinsic temporal structures and diverse complexity, would serve as a valuable evaluation probe for developing effective long-term video understanding systems in the future. Data and Zero-shot model evaluation code are open-sourced for both public and commercial use under the Ego4D license at http://egoschema.github.io

📄 PDF Abstract BibTeX arXiv:2308.09126

Code (1)

egoschema/egoschema 공식 구현 pytorch

Tasks

DiagnosticEgoSchemaFormMultiple-choiceQuestion AnsweringVideo Question AnsweringVideo UnderstandingZero-Shot Video Question Answer

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

HCQA @ Ego4D EgoSchema Challenge 2024

2024-06-22 · Haoyu Zhang, Yuquan Xie, Yisen Feng, Zaijing Li 외

In this report, we present our champion solution for Ego4D EgoSchema Challenge in CVPR 2024. To deeply integrate the powerful egocentric captioning model and question reasoning model, we propose a novel Hierarchical Comp…

Caption GenerationEgoSchemaMultiple-choice+2

From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

2026-08-31 · Can Zhang, Baofeng Zhang, Xiaotian Han, Junyuan Shang 외 arxiv

Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurrence coverage, or must discriminate competing hypotheses---which can …

LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos

2023-12-07 · Ying Wang, Yanlai Yang, Mengye Ren

In this paper we introduce LifelongMemory, a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval. LifelongMemory generates concise video activity…

EgoSchemaFormQuestion AnsweringRetrieval

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

2025-09-29 · Jiahua Li, Zhanhe Zhang, Chenghao Xu, Zhe Xu 외 arxiv

Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model (LLM)-based approaches have advanced lo…

DrVideo: Document Retrieval Based Long Video Understanding

2024-06-18 · CVPR 2025 1 · Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 외

Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos …

document understandingEgoSchemaMMERetrieval+2