paper-with-me

홈 › Papers

Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation

2026-02-12 · Shuo Lu, Jianjie Cheng, Yinuo Xu, Yongcan Yu, Lijun Sheng, Peijie Wang, Siru Jiang, Yongguan Hu, Run Ling, Yihua Shao, Ao Ma, Wei Feng, Lingxiao He, Meng Wang, Qianlong Xie, Xingxing Wang, Nicu Sebe, Ran He, Jian Liang arxiv

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the capacity to parse and manipulate two- and three-dimensional relations, remains unclear. Humans easily solve textbook-style spatial reasoning problems with over 95\% accuracy, but we find that most leading MLLMs fail to reach even 60\% on the same tasks. This striking gap highlights spatial reasoning as a fundamental weakness of current models. To investigate this gap, we present \emph{MathSpatial}, the first large-scale and systematic dataset resource dedicated to mathematical spatial reasoning in MLLMs. \emph{MathSpatial} provides two complementary subsets: (i)~\emph{MathSpatial-Bench}, a rigorously curated evaluation set of 2{,}000 problems spanning 3 categories and 11 subtypes, designed to isolate spatial reasoning from perceptual noise; and (ii)~\emph{MathSpatial-Corpus}, a training set of 8{,}000 problems equipped with verified solutions and structured reasoning traces. All problems are sourced from authentic educational materials and undergo multi-stage quality control including deduplication, geometric consistency checking, and cross-validated solution verification. Benchmarking 16 leading MLLMs on \emph{MathSpatial-Bench} reveals that spatial reasoning remains a fundamental bottleneck: even GPT-5 lags behind human performance by over 35 percentage points, with particularly poor results on abstract deduction tasks. We further show that training on \emph{MathSpatial-Corpus} yields consistent improvements across model families, demonstrating the dataset's practical value for advancing spatial reasoning capabilities. \emph{MathSpatial} is publicly available at https://shuolucs.github.io/MathSpatial.

📄 PDF Abstract BibTeX arXiv:2602.11635

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical ReasoningSpatial Reasoning

Similar Papers 제목 키워드 기반

Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist

2024-07-11 · ZiHao Zhou, Shudong Liu, Maizhen Ning, Wei Liu 외

Exceptional mathematical reasoning ability is one of the key features that demonstrate the power of large language models (LLMs). How to comprehensively define and evaluate the mathematical abilities of LLMs, and even re…

GSM8KMathMathematical Reasoning

ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints

2025-11-23 · Rui Xu, Dakuan Lu, Zicheng Zhao, Xiaoyu Tan 외 arxiv

Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimo…

Natural Language UnderstandingReinforcement LearningSpatial ReasoningCode Generation

Do MLLMs Really Understand the Charts?

2025-08-27 · Xiao Zhang, Dongyuan Li, Liuyu Xiang, Yao Zhang 외 arxiv

Although Multimodal Large Language Models (MLLMs) have demonstrated increasingly impressive performance in chart understanding, most of them exhibit alarming hallucinations and significant performance degradation when ha…

Visual Reasoning

Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?

2025-05-16 · Tairan Fu, Miguel González, Javier Conde, Elena Merino-Gómez 외

Multimodal Large Language Models which can answer complex questions on an image struggle to tell the time on analog clocks. This is probably due to the lack of images with clocks at different times in their training set.…

MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark

2024-08-14 · Minxuan Zhou, Hao Liang, Tianpeng Li, Zhiyu Wu 외

With the development of Multimodal Large Language Models (MLLMs), the evaluation of multimodal models in the context of mathematical problems has become a valuable research field. Multimodal visual-textual mathematical r…

MathMathematical Reasoning