paper-with-me

Papers

RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos

2025-10-10 · Zixi Yang, Jiapeng Li, Muxi Diao, Yinuo Jing, Kongming Liang arxiv

Recently, Multi-modal Large Language Models (MLLMs) have demonstrated significant performance across various video understanding tasks. However, their robustness, particularly when faced with manipulated video content, remains largely unexplored. In this paper, we introduce Ro-Bench, the first benchmark for evaluating MLLMs on dynamic out-of-distribution (OOD) counterfactual video test sets. Ro-Bench incorporates high-quality, diverse and temporally relevant video data, by editing Style, Object, Background and their compositions. We evaluated eight recent video MLLMs and found that current models exhibit substantial performance degradation on Ro-Bench when exposed to counterfactual video content. Furthermore, we demonstrate that fine-tuning MLLMs with counterfactual data enhances robustness, achieving a 21.73% performance increase on Ro-Bench and a 12.78% improvement across 20 tasks in the MVBench dataset. These findings underscore the effectiveness of counterfactual data in enhancing the video understanding ability of MLLMs. The code and data will be released shortly.

📄 PDF Abstract BibTeX arXiv:2510.08936

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input

2025-10-19 · Chenxu Li, Zhicai Wang, Yuan Sheng, Xingyu Zhu 외 arxiv

Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robust…

MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams

2025-08-09 · Pengfei Zhou, Xiaopeng Peng, Fanrui Zhang, Zhaopan Xu 외 arxiv

Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intel…

ChEF: A Comprehensive Evaluation Framework for Standardized Assessment of Multimodal Large Language Models

2023-11-05 · Zhelun Shi, Zhipin Wang, Hongxing Fan, Zhenfei Yin 외

Multimodal Large Language Models (MLLMs) have shown impressive abilities in interacting with visual content with myriad potential downstream tasks. However, even though a list of benchmarks has been proposed, the capabil…

HallucinationIn-Context LearningInstruction FollowingQuestion Answering

EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs

2025-11-23 · Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang 외 arxiv

Multimodal large language models (MLLMs) have made significant advancements in event-based vision, yet the comprehensive evaluation of their capabilities within a unified benchmark remains largely unexplored. In this wor…

Event-based visionSpatial Reasoning

MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness

2025-05-26 · Yunlong Tang, Pinxin Liu, Mingqian Feng, Zhangyun Tan 외

Understanding perspective is fundamental to human visual perception, yet the extent to which multimodal large language models (MLLMs) internalize perspective geometry remains unclear. We introduce MMPerspective, the firs…