paper-with-me

홈 › Papers

Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs

2025-05-28 · Aditya Kanade, Tanuja Ganu

Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a critical bottleneck. Strikingly, MLLMs can produce correct answers even while misinterpreting crucial visual elements, masking these underlying failures. Our preliminary study on a joint perception-reasoning dataset revealed that for one leading MLLM, 29% of its correct answers to reasoning questions still exhibited visual perception errors. To systematically address this, we introduce "Do You See Me", a scalable benchmark with 1,758 images and 2,612 questions. It spans seven human-psychology inspired subtasks in 2D and 3D, featuring controllable complexity to rigorously evaluate MLLM visual skills. Our findings on 3 leading closed-source and 5 major open-source models reveal a stark deficit: humans achieve 96.49% accuracy, while top MLLMs average below 50%. This performance gap widens rapidly with increased task complexity (e.g., from 12% to 45% in the visual form constancy subtask). Further analysis into the root causes suggests that failures stem from challenges like misallocated visual attention and the instability of internal representations for fine-grained details, especially at or below encoder patch resolution. This underscores an urgent need for MLLMs with truly robust visual perception. The benchmark dataset, source code and evaluation scripts are available at https://github.com/microsoft/Do-You-See-Me.

📄 PDF Abstract BibTeX arXiv:2506.02022

Code (1)

microsoft/do-you-see-me 공식 구현

Similar Papers 제목 키워드 기반

A new visual quality metric for Evaluating the performance of multidimensional projections

2024-07-23 · Maniru Ibrahim, Thales Vieira

Multidimensional projections (MP) are among the most essential approaches in the visual analysis of multidimensional data. It transforms multidimensional data into two-dimensional representations that may be shown as sca…

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

2026-07-27 · Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang 외 hf

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: hol…

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

2024-10-07 · Ziyue Wang, Chi Chen, Fuwen Luo, Yurui Dong 외

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimod…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models

2024-12-30 · Shangyu Xing, Changhao Xiang, Yuteng Han, Yifan Yue 외

Multimodal large language models (MLLMs) have achieved significant advancements in integrating visual and linguistic understanding. While existing benchmarks evaluate these models in context-rich, real-life scenarios, th…

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs

2026-02-20 · Ziqiao Shang, Lingyue Ge, Zi-Jian Cheng, Shi-Yu Tian 외 arxiv

Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for rigorously measuring their reasoning capa…

Multimodal Reasoning