paper-with-me

Papers

MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues

2025-10-20 · Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Wang, Yongqian Wen, Yuanxing Zhang, Haoxuan Hu, Zhiyu Pan, Yibing Huang, Zhidong Gan, Yonghong Lin, An Ping, Shihao Li, Yanghai Wang, Tianhao Peng, Jiaheng Liu arxiv

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. To bridge this gap, we introduce MT-Video-Bench, a holistic video understanding benchmark for evaluating MLLMs in multi-turn dialogues. Specifically, our MT-Video-Bench mainly assesses 6 core competencies that focus on perceptivity and interactivity, encompassing 1,000 meticulously curated multi-turn dialogues from diverse domains. These capabilities are rigorously aligned with real-world applications, such as interactive sports analysis and multi-turn video-based intelligent tutoring. With MT-Video-Bench, we extensively evaluate various state-of-the-art open-source and closed-source MLLMs, revealing their significant performance discrepancies and limitations in handling multi-turn video dialogues. The benchmark will be publicly available to foster future research.

📄 PDF Abstract BibTeX arXiv:2510.17722

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding

2025-03-31 · Qi Wu, Quanlong Zheng, Yanhao Zhang, Junlin Xie 외

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant…

Video Understanding

CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition

2025-03-30 · Jongseo Lee, Joohyun Chang, DongHo Lee, Jinwoo Choi

We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing …

Action ClassificationAction RecognitionAudio ClassificationVideo Recognition+1

MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence

2025-12-11 · Jingli Lin, Runsen Xu, Shaohao Zhu, Sihan Yang 외 arxiv

Spatial understanding over continuous visual input is crucial for MLLMs to evolve into general-purpose assistants in physical environments. Yet there is still no comprehensive benchmark that holistically assesses the pro…

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

2024-06-20 · Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao 외

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative…

FormVideo Understanding

ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

2024-12-12 · CVPR 2025 1 · Ali Athar, Xueqing Deng, Liang-Chieh Chen

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body…

Phrase GroundingQuestion AnsweringSegmentationSemantic Segmentation+2