paper-with-me

홈 › Papers

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

2025-10-07 · Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman arxiv

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application to video understanding remains underexplored. This paper presents a systematic analysis revealing that CoT often degrades performance in video reasoning, generating verbose but misleading internal monologues, and leading to hallucinated visual details and overridden correct intuitions - a phenomenon we term "visual thinking drift". We explain this drift through a Bayesian lens, positing that CoT traces often diverge from actual visual evidence, instead amplifying internal biases or language priors, causing models to storytell rather than engage in grounded reasoning. To counteract this, we introduce Visual Evidence Reward (VER), a novel reinforcement learning framework that explicitly rewards the generation of reasoning traces that are verifiably grounded in visual evidence. Comprehensive evaluation across 10 diverse video understanding benchmarks demonstrates that our Video-VER consistently achieves top performance. Our work sheds light on the distinct challenges of video-centric reasoning and encourages the development of AI that robustly grounds its inferences in visual evidence - for large multimodal models that not only "think before answering", but also "see while thinking".

📄 PDF Abstract BibTeX arXiv:2510.06077

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Beyond Uncertainty: Evidential Deep Learning for Robust Video Temporal Grounding

2024-08-29 · Kaijing Ma, Haojian Huang, Jin Chen, Haodong Chen 외

Existing Video Temporal Grounding (VTG) models excel in accuracy but often overlook open-world challenges posed by open-vocabulary queries and untrimmed videos. This leads to unreliable predictions for noisy, corrupted, …

cross-modal alignmentDeep Learning

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

2026-02-08 · Wenqi Liu, Yunxiao Wang, Shijie Ma, Meng Liu 외 arxiv

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-vid…

Reinforcement LearningQuestion AnsweringVideo Grounding

Ego-Grounding for Personalized Question-Answering in Egocentric Videos

2026-04-02 · Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela Yao arxiv

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this …

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

2025-08-21 · Pengcheng Fang, Yuxia Chen, Rui Guo arxiv

Understanding videos requires more than answering open ended questions, it demands the ability to pinpoint when events occur and how entities interact across time. While recent Video LLMs have achieved remarkable progres…

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation

2026-05-16 · Pengyu Yan, Akhil Gorugantu, Mahesh Bhosale, Abdul Wasi 외 arxiv

Multi-video event understanding demands models that can locate and attribute query-relevant evidence scattered across long, heterogeneous video corpora. Existing large vision-language models (LVLMs) often underperform in…

Object DetectionVisual Reasoning