paper-with-me

Papers

ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models

2025-12-01 · Xusen Hei, Jiali Chen, Jinyu Yang, Mengchen Zhao, Yi Cai arxiv

As multimodal large language models (MLLMs) frequently exhibit errors in complex video reasoning scenarios, correcting these errors is critical for uncovering their weaknesses and improving performance. However, existing benchmarks lack systematic evaluation of MLLMs' ability to identify and correct these video reasoning errors. To bridge this gap, we propose ViRectify, a comprehensive benchmark to evaluate their fine-grained correction capability. Through an AI-assisted annotation pipeline with human verification, we construct a dataset of over 30K instances spanning dynamic perception, scientific reasoning, and embodied decision-making domains. In ViRectify, we challenge MLLMs to perform step-wise error identification and generate rationales with key video evidence grounding. In addition, we further propose the trajectory evidence-driven correction framework, comprising step-wise error trajectory and reward modeling on visual evidence-grounded correction. It encourages the model to explicitly concentrate on error propagation and key timestamps for correction. Extensive evaluation across 16 advanced MLLMs demonstrates that our ViRectify serves as a challenging testbed, where GPT-5 achieves only 31.94% correction accuracy. Our framework enables a Qwen2.5-VL-7B to consistently outperform the variants of 72B on ViRectify, showing the effectiveness of our approach. Further analysis uncovers systematic asymmetries in error correction across models, and our dataset is also a valuable data resource to perform reflection learning. We believe ViRectify provides a new direction for comprehensively evaluating the advanced MLLMs in video reasoning.

📄 PDF Abstract BibTeX arXiv:2512.01424

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VGI-Bench: Probing Visual Intelligence in Video Generation Models

2026-08-20 · Xuan He, Cong Wei, Yuhao Cheng, Linrui Ma 외 arxiv

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned wi…

Visual ReasoningVideo Generation

Live Interactive Training for Video Segmentation

2026-03-27 · Xinyu Yang, Haozheng Yu, Yihong Sun, Bharath Hariharan 외 arxiv

Interactive video segmentation often requires many user interventions for robust performance in challenging scenarios (e.g., occlusions, object separations, camouflage, etc.). Yet, even state-of-the-art models like SAM2 …

Fine-Grained Image ClassificationVideo Segmentation

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

2025-05-19 · Ziyang Ma, Yinghao Ma, Yanqiao Zhu, Chen Yang 외

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-an…

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

2026-07-17 · Wei Feng, Xin Wang, Yu-Wei Zhan, Yuwei Zhou 외 arxiv

Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets a…

Reinforcement Learning

STAR: A Benchmark for Situated Reasoning in Real-World Videos

2024-05-15 · NeurIPS 2021 12 · Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum 외

Reasoning in the real world is not divorced from situations. How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence. This pa…

DiagnosticLogical ReasoningQuestion Answering