paper-with-me

Papers

V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction

2025-03-22 · Yiming Zhao, Yu Zeng, Yukun Qi, Yaoyang Liu, Lin Chen, Zehui Chen, Xikun Bao, Jie Zhao, Feng Zhao

Large Vision-Language Models (LVLMs) have made significant progress in the field of video understanding recently. However, current benchmarks uniformly lean on text prompts for evaluation, which often necessitate complex referential language and fail to provide precise spatial and temporal references. This limitation diminishes the experience and efficiency of human-model interaction. To address this limitation, we propose the Video Visual Prompt Benchmark(V2P-Bench), a comprehensive benchmark specifically designed to evaluate LVLMs' video understanding capabilities in multimodal human-model interaction scenarios. V2P-Bench includes 980 unique videos and 1,172 QA pairs, covering 5 main tasks and 12 dimensions, facilitating instance-level fine-grained understanding aligned with human cognition. Benchmarking results reveal that even the most powerful models perform poorly on V2P-Bench (65.4% for GPT-4o and 67.9% for Gemini-1.5-Pro), significantly lower than the human experts' 88.3%, highlighting the current shortcomings of LVLMs in understanding video visual prompts. We hope V2P-Bench will serve as a foundation for advancing multimodal human-model interaction and video understanding evaluation. Project page: https://github.com/gaotiexinqu/V2P-Bench.

📄 PDF Abstract BibTeX arXiv:2503.17736

Code (1)

gaotiexinqu/v2p-bench 공식 구현 pytorch

Tasks

BenchmarkingVideo Understanding

Similar Papers 제목 키워드 기반

Video-Oasis: Rethinking Evaluation of Video Understanding

2026-07-02 · Geuntaek Lim, Sungjune Park, Jaeyun Lee, Inwoong Lee 외 hf

The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have e…

MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues

2025-10-20 · Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Wang 외 arxiv

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question a…

Question Answering

TVBench: Redesigning Video-Language Evaluation

2024-10-10 · Daniel Cores, Michael Dorkenwald, Manuel Mucientes, Cees G. M. Snoek 외

Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating these video models presents its own unique challenges, for which se…

Multiple-choiceOpen-Ended Question AnsweringQuestion AnsweringVideo Understanding+2

DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM

2024-10-03 · Xuchen Li, Shiyu Hu, Xiaokun Feng, Dailing Zhang 외

Visual language tracking (VLT) has emerged as a cutting-edge research area, harnessing linguistic data to enhance algorithms with multi-modal inputs and broadening the scope of traditional single object tracking (SOT) to…

Object TrackingVideo Understanding

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

2026-08-26 · Hongbo Liu, Peixian Chen, Sihan Liu, Peiyuan Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires …

Instruction Following