paper-with-me

홈 › Papers

HoloCount: A Holistic Visual Counting Benchmark for MLLMs

2026-07-07 · Jinhong Deng, Limeng Qiao, Guanglu Wan arxiv

Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by persistent numerical hallucinations. Existing counting benchmarks primarily focus on basic perception in simplified contexts, failing to capture the complex failure modes that emerge under logical constraints or adversarial conditions. To address these limitations, we introduce HoloCount, a holistic and diagnostically rich benchmark structured around a three-level hierarchical taxonomy. HoloCount evaluates MLLMs across: (1) Semantic Counting, focusing on atomic and property-based enumeration; (2) Analytical Counting, assessing logical composition through spatial and set-based reasoning; and (3) Robustness Testing, probing model integrity against adverse scenarios and grounded counter-priors, such as high-density scenes and linguistic biases. Through an exhaustive evaluation of over 20 state-of-the-art MLLMs, we reveal a critical performance gap: even top-tier models degrade significantly as tasks transition from perception to complex analytical reasoning and adverse scenarios. Our findings provide a systematic landscape of current MLLM counting capabilities and offer a roadmap for developing more grounded and reliable multimodal systems. The dataset is available at https://mm-mvr.github.io/HoloCount/.

📄 PDF Abstract BibTeX arXiv:2607.06420

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingSpatial Reasoning

Similar Papers 제목 키워드 기반

ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams

2026-04-17 · Qiang Xu, Shengyuan Bai, Yu Wang, He Cao 외 arxiv

Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reasoning over simple linear diagrams. However, when faced with complex topological structures involving branching paths, conve…

Visual Reasoning

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

2025-06-05 · Lidong Lu, Guo Chen, Zhiqi Li, Yicheng Liu 외

Despite progress in video understanding, current MLLMs struggle with counting tasks. Existing benchmarks are limited by short videos, close-set queries, lack of clue annotations, and weak multimodal coverage. In this pap…

BenchmarkingVideo Understanding

TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs

2026-06-24 · Yu-Yang Chen, Lan-Zhe Guo arxiv

Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexity remains poorly understood. We introdu…

Visual Question AnsweringVisual ReasoningObject Counting

MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues

2025-10-20 · Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Wang 외 arxiv

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question a…

Question Answering

Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

2023-09-25 · HaoNing Wu, ZiCheng Zhang, Erli Zhang, Chaofeng Chen 외

The rapid evolution of Multi-modality Large Language Models (MLLMs) has catalyzed a shift in computer vision from specialized models to general-purpose foundation models. Nevertheless, there is still an inadequacy in ass…

Image Quality Assessment