paper-with-me

Papers

MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes

2025-08-24 · Nilay Pande, Sahiti Yerramilli, Jayant Sravan Tamarapalli, Rynaa Grover arxiv

A key frontier for Multimodal Large Language Models (MLLMs) is the ability to perform deep mathematical and spatial reasoning directly from images, moving beyond their established success in semantic description. Mathematical surface plots provide a rigorous testbed for this capability, as they isolate the task of reasoning from the semantic noise common in natural images. To measure progress on this frontier, we introduce MaRVL-QA (Mathematical Reasoning over Visual Landscapes), a new benchmark designed to quantitatively evaluate these core reasoning skills. The benchmark comprises two novel tasks: Topological Counting, identifying and enumerating features like local maxima; and Transformation Recognition, recognizing applied geometric transformations. Generated from a curated library of functions with rigorous ambiguity filtering, our evaluation on MaRVL-QA reveals that even state-of-the-art MLLMs struggle significantly, often resorting to superficial heuristics instead of robust spatial reasoning. MaRVL-QA provides a challenging new tool for the research community to measure progress, expose model limitations, and guide the development of MLLMs with more profound reasoning abilities.

📄 PDF Abstract BibTeX arXiv:2508.17180

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical ReasoningSpatial Reasoning

Similar Papers 제목 키워드 기반

Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing

2024-02-08 · Yong Cao, Wenyan Li, Jiaang Li, Yifei Yuan 외

Pretrained large Vision-Language models have drawn considerable interest in recent years due to their remarkable performance. Despite considerable efforts to assess these models from diverse perspectives, the extent of v…

Image CaptioningTAG

MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models

2026-01-28 · Xunlan Zhou, Xuanlin Chen, Shaowei Zhang, ShengHua Wan 외 arxiv

Designing dense reward functions is pivotal for efficient robotic Reinforcement Learning (RL). However, most dense rewards rely on manual engineering, which fundamentally limits the scalability and automation of reinforc…

Reinforcement Learning

MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts

2025-02-28 · CVPR 2025 1 · Peijie Wang, Zhong-Zhi Li, Fei Yin, Xin Yang 외

Multimodal Large Language Models (MLLMs) have shown promising capabilities in mathematical reasoning within visual contexts across various datasets. However, most existing multimodal math benchmarks are limited to single…

MathMathematical ReasoningMultiple-choice

MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?

2025-11-28 · Yuandong Wang, Yao Cui, Yuxin Zhao, Zhen Yang 외 arxiv

Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reasoning remains unclear. Existing benchmark…

Mathematical Reasoning

CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images

2025-10-13 · Chengqi Duan, Kaiyue Sun, Rongyao Fang, Manyuan Zhang 외 arxiv

Recent advances in Large Language Models (LLMs) and Vision Language Models (VLMs) have shown significant progress in mathematical reasoning, yet they still face a critical bottleneck with problems requiring visual assist…

Mathematical ReasoningVisual Reasoning