paper-with-me

홈 › Papers

KidVis: Do Multimodal Large Language Models Possess the Visual Perceptual Capabilities of a 6-Year-Old?

2026-01-13 · Xianfeng Wang, Kaiwei Zhang, Qi Jia, Zijian Chen, Guangtao Zhai, Xiongkuo Min arxiv

While Multimodal Large Language Models (MLLMs) have demonstrated impressive proficiency in high-level reasoning tasks, such as complex diagrammatic interpretation, it remains an open question whether they possess the fundamental visual primitives comparable to human intuition. To investigate this, we introduce KidVis, a novel benchmark grounded in the theory of human visual development. KidVis deconstructs visual intelligence into six atomic capabilities - Concentration, Tracking, Discrimination, Memory, Spatial, and Closure - already possessed by 6-7 year old children, comprising 10 categories of low-semantic-dependent visual tasks. Evaluating 20 state-of-the-art MLLMs against a human physiological baseline reveals a stark performance disparity. Results indicate that while human children achieve a near-perfect average score of 95.32, the state-of-the-art GPT-5 attains only 67.33. Crucially, we observe a "Scaling Law Paradox": simply increasing model parameters fails to yield linear improvements in these foundational visual capabilities. This study confirms that current MLLMs, despite their reasoning prowess, lack the essential physiological perceptual primitives required for generalized visual intelligence.

📄 PDF Abstract BibTeX arXiv:2601.08292

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

2025-06-03 · Yang Yao, Lingyu Li, Jiaxin Song, Chiyu Chen 외

As Multimodal Large Language Models (MLLMs) continue to evolve, their cognitive and reasoning capabilities have seen remarkable progress. However, challenges in visual fine-grained perception and commonsense causal infer…

Causal Inference

AIN: The Arabic INclusive Large Multimodal Model

2025-01-31 · Ahmed Heakl, Sara Ghaboura, Omkar Thawkar, Fahad Shahbaz Khan 외

Amid the swift progress of large language models (LLMs) and their evolution into large multimodal models (LMMs), significant strides have been made in high-resource languages such as English and Chinese. While Arabic LLM…

document understandingmodelVideo Understanding

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

2024-12-18 · CVPR 2025 1 · Jihan Yang, Shusheng Yang, Anjali W. Gupta, Rilyn Han 외

Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language Models (MLLMs) trained on million-scale video datasets also ``think in space''…

Question AnsweringSpatial Reasoning

VCoder: Versatile Vision Encoders for Multimodal Large Language Models

2023-12-21 · CVPR 2024 1 · Jitesh Jain, Jianwei Yang, Humphrey Shi

Humans possess the remarkable skill of Visual Perception, the ability to see and understand the seen, helping them make sense of the visual world and, in turn, reason. Multimodal Large Language Models (MLLM) have recentl…

Image CaptioningImage GenerationObjectQuestion Answering+2

Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!

2024-10-01 · Jiwan Chung, Seungwon Lim, Jaehyun Jeon, Seungbeen Lee 외

Humans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning. Faced with challenges like lexical ambiguity in text, we supplement this with other modalitie…