paper-with-me

Papers

GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models

2024-12-30 · Shangyu Xing, Changhao Xiang, Yuteng Han, Yifan Yue, Zhen Wu, Xinyu Liu, Zhangtai Wu, Fei Zhao, Xinyu Dai

Multimodal large language models (MLLMs) have achieved significant advancements in integrating visual and linguistic understanding. While existing benchmarks evaluate these models in context-rich, real-life scenarios, they often overlook fundamental perceptual skills essential for environments deviating from everyday realism. In particular, geometric perception, the ability to interpret spatial relationships and abstract visual patterns, remains underexplored. To address this limitation, we introduce GePBench, a novel benchmark designed to assess the geometric perception capabilities of MLLMs. Results from extensive evaluations reveal that current state-of-the-art MLLMs exhibit significant deficiencies in such tasks. Additionally, we demonstrate that models trained with data sourced from GePBench show notable improvements on a wide range of downstream tasks, underscoring the importance of geometric perception as a foundation for advanced multimodal applications. Our code and datasets will be publicly available.

📄 PDF Abstract BibTeX arXiv:2412.21036

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MATHGLANCE: Multimodal Large Language Models Do Not Know Where to Look in Mathematical Diagrams

2025-03-26 · Yanpeng Sun, Shan Zhang, Wei Tang, Aotian Chen 외

Diagrams serve as a fundamental form of visual language, representing complex concepts and their inter-relationships through structured symbols, shapes, and spatial arrangements. Unlike natural images, their inherently s…

Mathematical ReasoningObject Counting

VertiCue-Bench: Diagnosing Whether MLLMs Use Height Cues to Resolve 2D Ambiguity in Remote Sensing Natural Scenes

2026-05-25 · Jing Huang, Duanchu Wang, Junjie Yang, Zihang Cheng 외 arxiv

Multimodal Large Language Models (MLLMs) have recently shown promising progress in geospatial reasoning. However, existing remote sensing benchmarks remain largely 2D-centric, evaluating models primarily on optical appea…

Scene Understanding

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly

2026-04-10 · Zhi Jing, Jinbin Qiao, Ouyang Lu, Jicong Ao 외 arxiv

Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. Recent methods based on vision-language models (VLMs) largely rely on coar…

Spatial ReasoningPoint Clouds

An Eye for an AI: Evaluating GPT-4o's Visual Perception Skills and Geometric Reasoning Skills Using Computer Graphics Questions

2024-10-22 · Tony Haoran Feng, Paul Denny, Burkhard C. Wünsche, Andrew Luxton-Reilly 외

CG (Computer Graphics) is a popular field of CS (Computer Science), but many students find this topic difficult due to it requiring a large number of skills, such as mathematics, programming, geometric reasoning, and cre…

Large Language Model

Tangram: Benchmark for Evaluating Geometric Element Recognition in Large Multimodal Models

2024-08-25 · Chao Zhang, Jiamin Tang, Jing Xiao

Significant advancements in Large Multimodal Models (LMMs) have enabled them to tackle complex problems involving visual-mathematical reasoning. However, their ability to identify geometric elements remains underexplored…

Mathematical Reasoning