paper-with-me

홈 › Papers

Understanding Counting Mechanisms in Large Language and Vision-Language Models

2025-11-21 · Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari, Mobin Bagherian, Sadegh Mohammadian, Mohammad Izadi, Mahdieh Soleymani Baghshah arxiv

Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these foundation models represent and compute numerical information in counting tasks. We use controlled experiments with repeated textual and visual items and analyze counting in LLMs and LVLMs through a set of behavioral, observational, and causal mediation analyses. To this end, we design a specialized tool, CountScope, for the mechanistic interpretability of numerical content. Results show that individual tokens or visual features encode latent positional count information that can be extracted and transferred across contexts. Layerwise analyses reveal a progressive emergence of numerical representations, with lower layers encoding small counts and higher layers representing larger ones. We identify an internal counter mechanism that updates with each item, stored mainly in the final token or region. In LVLMs, numerical information also appears in visual embeddings, shifting between background and foreground regions depending on spatial composition. We further reveal that models rely on structural cues such as separators in text, which act as shortcuts for tracking item counts and strongly influence the accuracy of numerical predictions. Overall, counting emerges as a structured, layerwise process in LLMs and follows the same general pattern in LVLMs, shaped by the properties of the vision encoder.

📄 PDF Abstract BibTeX arXiv:2511.17699

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models

2026-03-19 · Liwei Che, Zhiyu Xue, Yihao Quan, Benlin Liu 외 arxiv

Counting serves as a simple but powerful test of a Large Vision-Language Model's (LVLM's) reasoning; it forces the model to identify each individual object and then add them all up. In this study, we investigate how LVLM…

Visual Reasoning

Object Counting with GPT-4o and GPT-5: A Comparative Study

2025-12-02 · Richard Füzesséry, Kaziwa Saleh, Sándor Szénási, Zoltán Vámossy arxiv

Zero-shot object counting attempts to estimate the number of object instances belonging to novel categories that the vision model performing the counting has never encountered during training. Existing methods typically …

Object Counting

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

2026-06-22 · Anindya Mondal, Sauradip Nag, Anjan Dutta arxiv

ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required. Our model is bu…

Object LocalizationImage GenerationObject CountingCrowd Counting

Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting

2025-10-06 · Xuyang Guo, Zekai Huang, Zhenmei Shi, Zhao Song 외 arxiv

Vision-Language Models (VLMs) have become a central focus of today's AI community, owing to their impressive abilities gained from training on large-scale vision-language data from the Web. These models have demonstrated…

Visual Reasoning

LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

2024-04-03 · Gabriela Ben Melech Stan, Estelle Aflalo, Raanan Yehezkel Rohekar, Anahita Bhiwandiwalla 외

In the rapidly evolving landscape of artificial intelligence, multi-modal large language models are emerging as a significant area of interest. These models, which combine various forms of data input, are becoming increa…

Language ModelingLanguage Modelling