paper-with-me

홈 › Papers

Can MLLMs Guide Me Home? A Benchmark Study on Fine-Grained Visual Reasoning from Transit Maps

2025-05-24 · Sicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong, Zikai Song, Jianke Zhu, Huan Wang, Xinchao Wang

Multimodal large language models (MLLMs) have recently achieved significant progress in visual tasks, including semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on complex tasks involving mathematics and logic. However, their capacity for reasoning tasks involving fine-grained visual understanding remains insufficiently evaluated. To address this gap, we introduce ReasonMap, a benchmark designed to assess the fine-grained visual understanding and spatial reasoning abilities of MLLMs. ReasonMap encompasses high-resolution transit maps from 30 cities across 13 countries and includes 1,008 question-answer pairs spanning two question types and three templates. Furthermore, we design a two-level evaluation pipeline that properly assesses answer correctness and quality. Comprehensive evaluations of 15 popular MLLMs, including both base and reasoning variants, reveal a counterintuitive pattern: among open-source models, base models outperform reasoning ones, while the opposite trend is observed in closed-source models. Additionally, performance generally degrades when visual inputs are masked, indicating that while MLLMs can leverage prior knowledge to answer some questions, fine-grained visual reasoning tasks still require genuine visual perception for strong performance. Our benchmark study offers new insights into visual reasoning and contributes to investigating the gap between open-source and closed-source models.

📄 PDF Abstract BibTeX arXiv:2505.18675

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingSpatial ReasoningVisual Reasoning

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models

2025-06-15 · Xinyi Zhao, Congjing Zhang, Pei Guo, Wei Li 외

Video anomaly detection (VAD) is essential for enhancing safety and security by identifying unusual events across different environments. Existing VAD benchmarks, however, are primarily designed for general-purpose scena…

Anomaly DetectionVideo Anomaly Detection

Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning

2025-04-01 · Ram Ramrakhya, Matthew Chang, Xavier Puig, Ruta Desai 외

Embodied agents operating in real-world environments must interpret ambiguous and under-specified human instructions. A capable household robot should recognize ambiguity and ask relevant clarification questions to infer…

Reinforcement Learning (RL)Vision-Language-Action

MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

2024-02-07 · Dongping Chen, Ruoxi Chen, Shilin Zhang, Yinuo Liu 외

Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence. However, assessing the utility of MLLMs presents considerable challeng…

MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?

2024-06-28 · Jinming Li, Yichen Zhu, Zhiyuan Xu, Jindong Gu 외

It is fundamentally challenging for robots to serve as useful assistants in human environments because this requires addressing a spectrum of sub-problems across robotics, including perception, language understanding, re…

Task PlanningVisual Reasoning

ShutterMuse: Capture-Time Photography Guidance with MLLMs

2026-06-24 · Jiayu Li, Yixiao Fang, Tianyu Hu, Wei Cheng 외 arxiv

Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evaluate post-hoc crop prediction and overlook subject-side recommendation…