paper-with-me

홈 › Papers

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos

2025-06-04 · Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. However, existing video benchmarks mainly focus on understanding tasks, which only require models to match frames mentioned in the question (hereafter referred to as "question frame") and perceive a few adjacent frames. To address this gap, we propose MMR-V: A Benchmark for Multimodal Deep Reasoning in Videos. The benchmark is characterized by the following features. (1) Long-range, multi-frame reasoning: Models are required to infer and analyze evidence frames that may be far from the question frame. (2) Beyond perception: Questions cannot be answered through direct perception alone but require reasoning over hidden information. (3) Reliability: All tasks are manually annotated, referencing extensive real-world user understanding to align with common perceptions. (4) Confusability: Carefully designed distractor annotation strategies to reduce model shortcuts. MMR-V consists of 317 videos and 1,257 tasks. Our experiments reveal that current models still struggle with multi-modal reasoning; even the best-performing model, o4-mini, achieves only 52.5% accuracy. Additionally, current reasoning enhancement strategies (Chain-of-Thought and scaling test-time compute) bring limited gains. Further analysis indicates that the CoT demanded for multi-modal reasoning differs from it in textual reasoning, which partly explains the limited performance gains. We hope that MMR-V can inspire further research into enhancing multi-modal reasoning capabilities.

📄 PDF Abstract BibTeX arXiv:2506.04141

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Debating the Unspoken: Role-Anchored Multi-Agent Reasoning for Half-Truth Detection

2026-04-21 · Yixuan Tang, Yirui Zhang, Hang Feng, Anthony K. H. Tung arxiv

Half-truths, claims that are factually correct yet misleading due to omitted context, remain a blind spot for fact verification systems focused on explicit falsehoods. Addressing such omission-based manipulation requires…

Fact Verification

What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models

2026-01-07 · Dasol Choi, Guijin Son, Hanwool Lee, Minhyuk Kim 외 arxiv

Current vision-language benchmarks predominantly feature well-structured questions with clear, explicit prompts. However, real user queries are often informal and underspecified. Users naturally leave much unsaid, relyin…

What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal News Previews

2026-01-09 · Fanxiao Li, Jiaying Wu, Tingchao Fu, Dayang Li 외 arxiv

Even when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial context, they lead readers to form judgments that diverge from what the full…

NOPE: A Corpus of Naturally-Occurring Presuppositions in English

2021-09-14 · CoNLL (EMNLP) 2021 11 · Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha 외

Understanding language requires grasping not only the overtly stated content, but also making inferences about things that were left unsaid. These inferences include presuppositions, a phenomenon by which a listener lear…

What Happens Next? Next Scene Prediction with a Unified Video Model

2025-12-15 · Xinjie Li, Zhimin Chen, Rui Zhao, Florian Schiffers 외 arxiv

Recent unified models for joint understanding and generation have significantly advanced visual generation capabilities. However, their focus on conventional tasks like text-to-video generation has left the temporal reas…

Text-to-Video GenerationReinforcement Learning