paper-with-me

Papers

How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

2024-05-06 · Muhammad Uzair Khattak, Muhammad Ferjad Naeem, Jameel Hassan, Muzammal Naseer, Federico Tombari, Fahad Shahbaz Khan, Salman Khan

Recent advancements in Large Language Models (LLMs) have led to the development of Video Large Multi-modal Models (Video-LMMs) that can handle a wide range of video understanding tasks. These models have the potential to be deployed in real-world applications such as robotics, AI assistants, medical surgery, and autonomous vehicles. The widespread adoption of Video-LMMs in our daily lives underscores the importance of ensuring and evaluating their robust performance in mirroring human-like reasoning and interaction capabilities in complex, real-world contexts. However, existing benchmarks for Video-LMMs primarily focus on general video comprehension abilities and neglect assessing their reasoning capabilities over complex videos in the real-world context, and robustness of these models through the lens of user prompts as text queries. In this paper, we present the Complex Video Reasoning and Robustness Evaluation Suite (CVRR-ES), a novel benchmark that comprehensively assesses the performance of Video-LMMs across 11 diverse real-world video dimensions. We evaluate 9 recent models, including both open-source and closed-source variants, and find that most of the Video-LMMs, especially open-source ones, struggle with robustness and reasoning when dealing with complex videos. Based on our analysis, we develop a training-free Dual-Step Contextual Prompting (DSCP) technique to enhance the performance of existing Video-LMMs. Our findings provide valuable insights for building the next generation of human-centric AI systems with advanced robustness and reasoning capabilities. Our dataset and code are publicly available at: https://mbzuai-oryx.github.io/CVRR-Evaluation-Suite/.

📄 PDF Abstract BibTeX arXiv:2405.03690

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous VehiclesVideo Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

2025-07-20 · Yuanhan Zhang, Yunice Chew, Yuhao Dong, Aria Leo 외 arxiv

Human intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintai…

Team of One: Cracking Complex Video QA with Model Synergy

2025-07-18 · Jun Xie, Zhaoran Zhao, Xiongjun Guan, Yingjian Zhu 외 arxiv

We propose a novel framework for open-ended video question answering that enhances reasoning depth and robustness in complex real-world scenarios, as benchmarked on the CVRR-ES dataset. Existing Video-Large Multimodal Mo…

Video Question AnsweringMultimodal Reasoning

Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

2023-05-23 · Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He 외

Despite exciting recent results showing vision-language systems' capacity to reason about images using natural language, their capacity for video reasoning remains under-explored. We motivate framing video reasoning as t…

DescriptiveVideo Prediction

MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning

2025-09-25 · Sicheng Tao, Jungang Li, Yibo Yan, Junyan Zhang 외 arxiv

Video reasoning has emerged as a critical capability for multimodal large language models (MLLMs), requiring models to move beyond static perception toward coherent understanding of temporal dynamics in complex scenes. Y…

Reinforcement Learning

Video-KTR: Reinforcing Video Reasoning via Key Token Attribution

2026-01-27 · Ziyue Wang, Sheng Jin, Zhongrong Zuo, Jiawei Wu 외 arxiv

Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models, yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token …

Reinforcement Learning