paper-with-me

홈 › Papers

Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs

2025-01-01 · CVPR 2025 1 · ZiCheng Zhang, Ziheng Jia, HaoNing Wu, Chunyi Li, Zijian Chen, Yingjie Zhou, Wei Sun, Xiaohong Liu, Xiongkuo Min, Weisi Lin, Guangtao Zhai

With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video quality understanding. To address this oversight, we introduce Q-Bench-Video in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. a) To ensure video source diversity, Q-Bench-Video encompasses videos from natural scenes, AI-generated Content (AIGC), and Computer Graphics (CG). b) Building on the traditional multiple-choice questions format with the Yes-or-No and What-How categories, we include Open-ended questions to better evaluate complex scenarios. Additionally, we incorporate the video pair quality comparison question to enhance comprehensiveness. c) Beyond the traditional Technical, Aesthetic, and Temporal distortions, we have expanded our evaluation aspects to include the dimension of AIGC distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to human performance. Through Q-Bench-Video, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choiceVideo GenerationVideo Understanding

Similar Papers 제목 키워드 기반

Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs

2024-09-30 · ZiCheng Zhang, Ziheng Jia, HaoNing Wu, Chunyi Li 외

With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the systematic exploration into video qua…

BenchmarkingMultiple-choiceVideo GenerationVideo Understanding

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

2026-05-08 · Yuancheng Wei, Linli Yao, Lei Li, Haojie Zhang 외 arxiv

Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality pr…

EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval

2024-07-23 · Thomas Hummel, Shyamgopal Karthik, Mariana-Iuliana Georgescu, Zeynep Akata

In Composed Video Retrieval, a video and a textual description which modifies the video content are provided as inputs to the model. The aim is to retrieve the relevant video with the modified content from a database of …

Re-RankingRetrievalVideo RetrievalVideo Understanding

ALLVB: All-in-One Long Video Understanding Benchmark

2025-03-10 · Xichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu 외

From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively short, which makes them inadequate for effec…

AllVideo Understanding

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

2024-12-31 · CVPR 2025 1 · Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng 외

Video Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spati…

ObjectVideo Understanding