paper-with-me

Papers

GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning

2023-12-19 · Mehran Kazemi, Hamidreza Alvari, Ankit Anand, Jialin Wu, Xi Chen, Radu Soricut

Large language models have shown impressive results for multi-hop mathematical reasoning when the input question is only textual. Many mathematical reasoning problems, however, contain both text and image. With the ever-increasing adoption of vision language models (VLMs), understanding their reasoning abilities for such problems is crucial. In this paper, we evaluate the reasoning capabilities of VLMs along various axes through the lens of geometry problems. We procedurally create a synthetic dataset of geometry questions with controllable difficulty levels along multiple axes, thus enabling a systematic evaluation. The empirical results obtained using our benchmark for state-of-the-art VLMs indicate that these models are not as capable in subjects like geometry (and, by generalization, other topics requiring similar reasoning) as suggested by previous benchmarks. This is made especially clear by the construction of our benchmark at various depth levels, since solving higher-depth problems requires long chains of reasoning rather than additional memorized knowledge. We release the dataset for further research in this area.

📄 PDF Abstract BibTeX arXiv:2312.12241

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models

2024-10-17 · Aditya Sharma, Aman Dalmia, Mehran Kazemi, Amal Zouaq 외

Geometry problem-solving demands advanced reasoning abilities to process multimodal inputs and employ mathematical knowledge effectively. Vision-language models (VLMs) have made significant progress in various multimodal…

Geometry Problem SolvingRAG

GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning

2025-04-17 · Liangyu Xu, Yingxiu Zhao, Jingyun Wang, Yingyao Wang 외

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit s…

Geometry Problem SolvingMultimodal Reasoning

CapGeo: A Caption-Assisted Approach to Geometric Reasoning

2025-10-10 · Yuying Li, Siyi Qian, Hao Liang, Leqi Zheng 외 arxiv

Geometric reasoning remains a core challenge for Multimodal Large Language Models (MLLMs). Even the most advanced closed-source systems, such as GPT-O3 and Gemini-2.5-Pro, still struggle to solve geometry problems reliab…

Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance

2026-04-02 · Jason Qiu, Zachary Meurer, Xavier Thomas, Deepti Ghadiyaram arxiv

This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks such as recognizing objects in canonical…

Spatial Reasoning

GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation

2025-12-30 · Yuan Feng, Yue Yang, Xiaohan He, Jiatong Zhao 외 arxiv

Geometric problem solving constitutes a critical branch of mathematical reasoning, requiring precise analysis of shapes and spatial relationships. Current evaluations of geometric reasoning in vision-language models (VLM…

Mathematical ReasoningAttribute Extraction