paper-with-me

Papers

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

2026-07-27 · Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou hf

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.

📄 PDF Abstract BibTeX arXiv:2607.24957

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models

2025-05-26 · Hyunsik Chae, Seungwoo Yoon, Jaden Park, Chloe Yewon Chun 외

Recent Vision-Language Models (VLMs) have demonstrated impressive multimodal comprehension and reasoning capabilities, yet they often struggle with trivially simple visual tasks. In this work, we focus on the domain of b…

VCoder: Versatile Vision Encoders for Multimodal Large Language Models

2023-12-21 · CVPR 2024 1 · Jitesh Jain, Jianwei Yang, Humphrey Shi

Humans possess the remarkable skill of Visual Perception, the ability to see and understand the seen, helping them make sense of the visual world and, in turn, reason. Multimodal Large Language Models (MLLM) have recentl…

Image CaptioningImage GenerationObjectQuestion Answering+2

Unveiling and Bridging the Functional Perception Gap in MLLMs: Atomic Visual Alignment and Hierarchical Evaluation via PET-Bench

2026-01-06 · Zanting Ye, Xiaolong Niu, Xuanbin Wu, Xu Han 외 arxiv

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in tasks such as abnormality detection and report generation for anatomical modalities, their capability in functional imaging remai…

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

2024-10-07 · Ziyue Wang, Chi Chen, Fuwen Luo, Yurui Dong 외

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimod…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis

2025-04-16 · Shravan Chaudhari, Trilokya Akula, Yoon Kim, Tom Blake

In this paper, we advance the study of AI-augmented reasoning in the context of Human-Computer Interaction (HCI), psychology and cognitive science, focusing on the critical task of visual perception. Specifically, we inv…