paper-with-me

홈 › Papers

VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

2025-03-19 · Tengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong Ming

Can Multimodal Large Language Models (MLLMs) develop an intuitive number sense similar to humans? Targeting this problem, we introduce Visual Number Benchmark (VisNumBench) to evaluate the number sense abilities of MLLMs across a wide range of visual numerical tasks. VisNumBench consists of about 1,900 multiple-choice question-answer pairs derived from both synthetic and real-world visual data, covering seven visual numerical attributes and four types of visual numerical estimation tasks. Our experiments on VisNumBench led to the following key findings: (i) The 17 MLLMs we tested, including open-source models such as Qwen2.5-VL and InternVL2.5, as well as proprietary models like GPT-4o and Gemini 2.0 Flash, perform significantly below human levels in number sense-related tasks. (ii) Multimodal mathematical models and multimodal chain-of-thought (CoT) models did not exhibit significant improvements in number sense abilities. (iii) Stronger MLLMs with larger parameter sizes and broader general abilities demonstrate modest gains in number sense abilities. We believe VisNumBench will serve as a valuable resource for the research community, encouraging further advancements in enhancing MLLMs' number sense abilities. All benchmark resources, including code and datasets, will be publicly available at https://wwwtttjjj.github.io/VisNumBench/.

📄 PDF Abstract BibTeX arXiv:2503.14939

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choice

Similar Papers 제목 키워드 기반

Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd

2026-02-02 · Yejin Son, Saejin Kim, Dongjun Min, Younjae Yu arxiv

Commonsense reasoning in multimodal contexts remains a foundational challenge in artificial intelligence. We introduce Multimodal UNcommonsense(MUN), a benchmark designed to evaluate models' ability to handle scenarios t…

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor

2024-12-08 · Jiali Chen, Xusen Hei, Yuqi Xue, Yuancheng Wei 외

Large multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image. However, the a…

MisconceptionsMultiple-choiceVisual Commonsense Reasoning

SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation

2026-04-02 · Haomin Zhuang, Xiangqi Wang, Yili Shen, Ying Cheng 외 arxiv

Large language models often default to step-by-step computation even when efficient numerical shortcuts are available. This raises a basic question: do they exhibit number sense in a human-like behavioral sense, i.e., th…

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

2025-02-06 · Jack Hong, Shilin Yan, Jiayin Cai, XiaoLong Jiang 외

In this paper, we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSens…

Video Understanding

GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning

2025-04-17 · Liangyu Xu, Yingxiu Zhao, Jingyun Wang, Yingyao Wang 외

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit s…

Geometry Problem SolvingMultimodal Reasoning