paper-with-me

홈 › Papers

GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory

2025-11-15 · Jeong Hun Yeo, Sangyun Chung, Sungjune Park, Dae Hoe Kim, Jinyoung Moon, Yong Man Ro arxiv

Long-video understanding remains a significant challenge for Multimodal Large Language Models (MLLMs) due to inherent token limitations and the complexity of capturing long-term temporal dependencies. Existing methods often fail to capture the global context and complex event relationships necessary for deep video reasoning. To address this, we introduce GCAgent, a novel Global-Context-Aware Agent framework that achieves comprehensive long-video understanding. Our core innovation is the Schematic and Narrative Episodic Memory. This memory structurally models events and their causal and temporal relations into a concise, organized context, fundamentally resolving the long-term dependency problem. Operating in a multi-stage Perception-Action-Reflection cycle, our GCAgent utilizes a Memory Manager to retrieve relevant episodic context for robust, context-aware inference. Extensive experiments confirm that GCAgent significantly enhances long-video understanding, achieving up to 23.5\% accuracy improvement on the Video-MME Long split over a strong MLLM baseline. Furthermore, our framework establishes state-of-the-art performance among comparable 7B-scale MLLMs, achieving 73.4\% accuracy on the Long split and the highest overall average (71.9\%) on the Video-MME benchmark, validating our agent-based reasoning paradigm and structured memory for cognitively-inspired long-video understanding.

📄 PDF Abstract BibTeX arXiv:2511.12027

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

2026-08-13 · Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi 외 arxiv

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks r…

Speaker Verification

NEST: Narrative Event Structures in Time for Long Video Understanding

2026-06-18 · Ali Asgarov, Kaushik Narasimhan, Najibul Haque Sarker, Hani Alomari 외 arxiv

Recent progress in vision-language models has enabled the processing of increasingly long video sequences, but the ability to handle extended token streams does not translate to understanding of narrative structure in lo…

Relation Extraction

Using Functional Schemas to Understand Social Media Narratives

2019-08-01 · WS 2019 8 · Xinru Yan, Aakanksha Naik, Yohan Jo, Carolyn Rose

We propose a novel take on understanding narratives in social media, focusing on learning {''}functional story schemas{''}, which consist of sets of stereotypical functional structures. We develop an unsupervised pipelin…

ClassificationGeneral Classificationtext-classificationText Classification

How does longer temporal context enhance multimodal narrative video processing in the brain?

2026-02-07 · Prachi Jindal, Anant Khandelwal, Manish Gupta, Bapi S. Raju 외 arxiv

Understanding how humans and artificial intelligence systems process complex narrative videos is a fundamental challenge at the intersection of neuroscience and machine learning. This study investigates how the temporal …

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding

2025-04-30 · CVPR 2025 1 · Chenkai Zhang, Yiming Lei, Zeming Liu, Haitao Leng 외

With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmark…

Video Understanding