paper-with-me

홈 › Papers

VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

2026-04-02 · Jiahao Meng, Tan Yue, Qi Xu, Haochen Wang, Zhongwei Ren, Weisong Liu, Yuhao Wang, Renrui Zhang, Yunhai Tong, Haodong Duan arxiv

Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual understanding and reasoning, and (2) answer correctness is often measured without verifying whether models identify the precise spatio-temporal evidence supporting their predictions. To address this, we present VideoZeroBench, a hierarchical benchmark designed for challenging long-video question answering that rigorously verifies spatio-temporal evidence. It comprises 500 manually annotated questions across 13 domains, paired with temporal intervals and spatial bounding boxes as evidence. To disentangle answering generation, temporal grounding, and spatial grounding, we introduce a five-level evaluation protocol that progressively tightens evidence requirements. Experiments show that even Gemini-3-Pro correctly answers fewer than 17% of questions under the standard end-to-end QA setting (Level-3). When grounding constraints are imposed, performance drops sharply: No model exceeds 1% accuracy when both correct answering and accurate spatio-temporal localization are required (Level-5), with most failing to achieve any correct grounded predictions. These results expose a significant gap between surface-level answer correctness and genuine evidence-based reasoning, revealing that grounded video understanding remains a bottleneck for long-video QA. We further analyze performance across minimal evidence spans, atomic abilities, and inference paradigms, providing insights for future research in grounded video reasoning. The benchmark and code will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2604.01569

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

2026-05-21 · Jinho Park, Youbin Kim, Hogun Park, Eunbyung Park arxiv

Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisely has become an essential challenge. However, existing spatio-tempor…

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

2025-10-27 · Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He 외 arxiv

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core chall…

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

2025-09-02 · Mennatullah Siam arxiv

Multi-modal large language models (MLLMs) have shown impressive generalization across tasks using images and text modalities. While their extension to video has enabled tasks such as video question answering and video ca…

Video Question AnsweringReferring ExpressionVideo CaptioningVisual Grounding

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

2025-04-09 · Xinhao Li, Ziang Yan, Desen Meng, Lu Dong 외

Recent advancements in reinforcement learning have significantly advanced the reasoning capabilities of multimodal large language models (MLLMs). While approaches such as Group Relative Policy Optimization (GRPO) and rul…

MVBenchObject TrackingVideo Understanding

Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph

2025-10-13 · Wentao Wang, Heqing Zou, Tianze Luo, Rui Huang 외 arxiv

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understanding. Existing spatio-temporal methods pr…

Reinforcement Learning