paper-with-me

홈 › Papers

VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models

2025-11-14 · Mingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li, Yue Qiu, Zekang Du, Mengyang Wu, Pingping Zhang, Kun Li, Hongzheng Yang, Wenao Ma, Jiaheng Wei, Qinbin Li, Kangcheng Liu, Wenqiang Lei arxiv

Multimodal large language models (MLLMs) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding. When querying specific regions or objects in an image, human users naturally use "visual prompts" (VPs), such as bounding boxes, to provide reference. However, no existing benchmark systematically evaluates the ability of MLLMs to interpret such VPs. This gap leaves it unclear whether current MLLMs can effectively recognize VPs, an intuitive prompting method for humans, and use them to solve problems. To address this limitation, we introduce VP-Bench, a benchmark for assessing MLLMs' capability in VP perception and utilization. VP-Bench employs a two-stage evaluation framework: Stage 1 examines models' ability to perceive VPs in natural scenes, using 30k visualized prompts spanning eight shapes and 355 attribute combinations. Stage 2 investigates the impact of VPs on downstream tasks, measuring their effectiveness in real-world problem-solving scenarios. Using VP-Bench, we evaluate 28 MLLMs, including proprietary systems (e.g., GPT-4o) and open-source models (e.g., InternVL3 and Qwen2.5-VL), and provide a comprehensive analysis of factors that affect VP understanding, such as variations in VP attributes, question arrangement, and model scale. VP-Bench establishes a new reference framework for studying how MLLMs comprehend and resolve grounded referring questions.

📄 PDF Abstract BibTeX arXiv:2511.11438

Code (0)

등록된 구현이 없습니다.

Tasks

Object Recognition

Similar Papers 제목 키워드 기반

Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

2024-03-29 · Weifeng Lin, Xinyu Wei, Ruichuan An, Peng Gao 외

The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level compr…

Instruction FollowingLanguage ModellingLarge Language Modelmultimodal interaction+4

VisFactor: Benchmarking Fundamental Visual Cognition in Multimodal Large Language Models

2025-02-23 · Jen-tse Huang, Dasen Dai, Jen-Yuan Huang, Youliang Yuan 외

Multimodal Large Language Models (MLLMs) have demonstrated remarkable advancements in multimodal understanding; however, their fundamental visual cognitive abilities remain largely underexplored. To bridge this gap, we i…

BenchmarkingSpatial ReasoningVisual Reasoning

VRPTEST: Evaluating Visual Referring Prompting in Large Multimodal Models

2023-12-07 · Zongjie Li, Chaozheng Wang, Chaowei Liu, Pingchuan Ma 외

With recent advancements in Large Multimodal Models (LMMs) across various domains, a novel prompting method called visual referring prompting has emerged, showing significant potential in enhancing human-computer interac…

ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts

2023-12-01 · CVPR 2024 1 · Mu Cai, Haotian Liu, Dennis Park, Siva Karthik Mustikovela 외

While existing large vision-language multimodal models focus on whole image understanding, there is a prominent gap in achieving region-specific comprehension. Current approaches that use textual coordinates or spatial e…

Visual Commonsense ReasoningVisual Prompting

Assessing the Value of Visual Input: A Benchmark of Multimodal Large Language Models for Robotic Path Planning

2025-07-16 · Jacinto Colan, Ana Davila, Yasuhisa Hasegawa arxiv

Large Language Models (LLMs) show potential for enhancing robotic path planning. This paper assesses visual input's utility for multimodal LLMs in such tasks via a comprehensive benchmark. We evaluated 15 multimodal LLMs…

Spatial Reasoning