paper-with-me

Papers

Narrative Aligned Long Form Video Question Answering

2026-03-19 · Rahul Jain, Keval Doshi, Burak Uzkent, Garin Kessler arxiv

Recent progress in multimodal large language models (MLLMs) has led to a surge of benchmarks for long-video reasoning. However, most existing benchmarks rely on localized cues and fail to capture narrative reasoning, the ability to track intentions, connect distant events, and reconstruct causal chains across an entire movie. We introduce NA-VQA, a benchmark designed to evaluate deep temporal and narrative reasoning in long-form videos. NA-VQA contains 88 full-length movies and 4.4K open-ended question-answer pairs, each grounded in multiple evidence spans labeled as Short, Medium, or Far to assess long-range dependencies. By requiring generative, multi-scene answers, NA-VQA tests whether models can integrate dispersed narrative information rather than rely on shallow pattern matching. To address the limitations of existing approaches, we propose Video-NaRA, a narrative-centric framework that builds event-level chains and stores them in a structured memory for retrieval during reasoning. Extensive experiments show that state-of-the-art MLLMs perform poorly on questions requiring far-range evidence, highlighting the need for explicit narrative modeling. Video-NaRA improves long-range reasoning performance by up to 3 percent, demonstrating its effectiveness in handling complex narrative structures. We will release NA-VQA upon publication.

📄 PDF Abstract BibTeX arXiv:2603.19481

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

Long Story Short: a Summarize-then-Search Method for Long Video Question Answering

2023-11-02 · Jiwan Chung, Youngjae Yu

Large language models such as GPT-3 have demonstrated an impressive capability to adapt to new tasks without requiring task-specific training data. This capability has been particularly effective in settings such as narr…

DiversityQuestion AnsweringVideo Question AnsweringVideo Question Answering (Level 3)+2

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

2026-08-13 · Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi 외 arxiv

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks r…

Speaker Verification

VideoAuteur: Towards Long Narrative Video Generation

2025-01-10 · Junfei Xiao, Feng Cheng, Lu Qi, Liangke Gui 외

Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informat…

Video Generation

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation

2025-07-15 · X. Feng, H. Yu, M. Wu, S. Hu 외 arxiv

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the g…

Question GenerationVideo Generation

SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation

2026-07-17 · Yilai Liu, Xin Zhang, Shiyuan Zhang, Hongyang Du arxiv

Maintaining recurring character identities across scene transitions and long temporal gaps is a central challenge in narrative long video generation. Methods targeting global consistency often retrieve memory using cues …

Video Generation