paper-with-me

Papers

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs

2025-04-21 · David Ma, Yuanxing Zhang, Jincheng Ren, Jarvis Guo, Yifan Yao, Zhenlin Wei, Zhenzhu Yang, Zhongyuan Peng, Boyu Feng, Jun Ma, Xiao Gu, Zhoufutu Wen, King Zhu, Yancheng He, Meng Cao, Shiwen Ni, Jiaheng Liu, Wenhao Huang, Ge Zhang, Xiaojie Jin

Existing evaluation frameworks for Multimodal Large Language Models (MLLMs) primarily focus on image reasoning or general video understanding tasks, largely overlooking the significant role of image context in video comprehension. To bridge this gap, we propose IV-Bench, the first comprehensive benchmark for evaluating Image-Grounded Video Perception and Reasoning. IV-Bench consists of 967 videos paired with 2,585 meticulously annotated image-text queries across 13 tasks (7 perception and 6 reasoning tasks) and 5 representative categories. Extensive evaluations of state-of-the-art open-source (e.g., InternVL2.5, Qwen2.5-VL) and closed-source (e.g., GPT-4o, Gemini2-Flash and Gemini2-Pro) MLLMs demonstrate that current models substantially underperform in image-grounded video Perception and Reasoning, merely achieving at most 28.9% accuracy. Further analysis reveals key factors influencing model performance on IV-Bench, including inference pattern, frame number, and resolution. Additionally, through a simple data synthesis approach, we demonstratethe challenges of IV- Bench extend beyond merely aligning the data format in the training proecss. These findings collectively provide valuable insights for future research. Our codes and data are released in https://github.com/multimodal-art-projection/IV-Bench.

📄 PDF Abstract BibTeX arXiv:2504.15415

Code (1)

multimodal-art-projection/iv-bench 공식 구현 pytorch

Tasks

Video Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

2026-07-17 · Runmao Yao, Kairui Hu, Yukang Cao, Ruisi Wang 외 hf

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without…

Video Generation

X2SAM: Any Segmentation in Images and Videos

2026-04-27 · Hao Wang, Limeng Qiao, Chi Zhang, Lin Ma 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos remains limited. Foundation segmentation mo…

Image SegmentationVideo Segmentation

Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark

2024-11-29 · Joseph Heyward, João Carreira, Dima Damen, Andrew Zisserman 외

Following the successful 2023 edition, we organised the Second Perception Test challenge as a half-day workshop alongside the IEEE/CVF European Conference on Computer Vision (ECCV) 2024, with the goal of benchmarking sta…

BenchmarkingGrounded Video Question AnsweringMultiple-choiceObject Tracking+5

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

2026-05-04 · Le Zhang, Jihan Yang, Soundarya Krishnan, Jimit Majmudar 외 arxiv

Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they are for. While existing benchmarks effectively evaluate the geometric …

Relational ReasoningSpatial Reasoning

Perception Test 2023: A Summary of the First Challenge And Outcome

2023-12-20 · Joseph Heyward, João Carreira, Dima Damen, Andrew Zisserman 외

The First Perception Test challenge was held as a half-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2023, with the goal of benchmarking state-of-the-art video models on the recen…

BenchmarkingGrounded Video Question AnsweringMultiple-choiceObject Tracking+3