paper-with-me

Papers

StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning

2026-04-25 · Xuanyue Zhong, Yuqiang Xie, Guanqun Bi, Jiangping Yang, Guibin Chen arxiv

Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see \textit{what is happening} but fail to reason \textit{why it matters}. This semantic gap stems from the lack of \textbf{Theory of Mind (ToM)}: the cognitive ability to infer implicit intentions, mental states, and narrative causality from surface-level observations. We introduce \textbf{StoryTR}, the first video moment retrieval benchmark requiring ToM reasoning, comprising 8.1k samples from narrative short-form videos (shorts/reels). These videos present an ideal testbed. Their high information density encodes meaning through subtle multimodal cues. For instance, a glance paired with a sigh carries entirely different semantics than the glance alone. Yet multimodal perception alone is insufficient; ToM is required to decode that a character `smiling'' may actually be `concealing hostility.'' To teach models this reasoning capability, we propose an \textbf{Agentic Data Pipeline} that generates training data with explicit three-tier ToM chains (intent decoding, narrative reasoning, boundary localization). Experiments reveal the severity of the reasoning gap: Gemini-3.0-Pro achieves only 0.53 Avg IoU on StoryTR. However, our 7B \textbf{Shorts-Moment} model, trained on ToM-guided data, improves +15.1\% relative IoU over baselines, demonstrating that \textit{narrative reasoning capability matters more than parameter scale}.

📄 PDF Abstract BibTeX arXiv:2604.23198

Code (0)

등록된 구현이 없습니다.

Tasks

Moment Retrieval

Similar Papers 제목 키워드 기반

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

2026-01-03 · Hyeonjeong Ha, Jinjin Ge, Bo Feng, Kaixin Ma 외 arxiv

Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored. True narrative und…

Narrative Aligned Long Form Video Question Answering

2026-03-19 · Rahul Jain, Keval Doshi, Burak Uzkent, Garin Kessler arxiv

Recent progress in multimodal large language models (MLLMs) has led to a surge of benchmarks for long-video reasoning. However, most existing benchmarks rely on localized cues and fail to capture narrative reasoning, the…

Video Question Answering

VideoMemory: Toward Consistent Video Generation via Memory Integration

2026-01-07 · Jinsong Zhou, Yihua Du, Xinli Xu, Luozhou Wang 외 arxiv

Maintaining consistent characters, props, and environments across multiple shots is a central challenge in narrative video generation. Existing models can produce high-quality short clips but often fail to preserve entit…

Video Generation

Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding

2026-06-04 · Qiuyu Tian, Fengyi Chen, Yiding Li, Youyong Kong 외 arxiv

Long-form narrative QA requires reasoning over evolving story worlds rather than isolated passages: answers may depend on earlier goals, changing character states, social relations, causal triggers, temporal position, an…

Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

2024-08-07 · Zi-Yi Dou, Xitong Yang, Tushar Nagarajan, Huiyu Wang 외

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse acti…

Multi-Instance RetrievalRepresentation LearningStyle Transfer