paper-with-me

Papers

GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation

2025-12-19 · Rang Li, Lei Li, Shuhuai Ren, Hao Tian, Shuhao Gu, Shicheng Li, Zihao Yue, Yudong Wang, Wenhan Ma, Zhe Yang, Jingyuan Ma, Zhifang Sui, Fuli Luo arxiv

Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing benchmarks, a fundamental question remains: can MLLMs truly visually ground with human-like sophistication, or are they merely pattern-matching on simplified datasets? Current benchmarks fail to capture real-world complexity where humans effortlessly navigate intricate references and recognize when grounding is impossible. To rigorously assess MLLMs' true capabilities, we introduce GroundingME, a benchmark that systematically challenges models across four critical dimensions: (1) Discriminative: distinguishing highly similar objects, (2) Spatial: understanding complex relational descriptions, (3) Limited: handling occlusions or tiny objects, and (4) Rejection: recognizing ungroundable queries. Through careful curation combining automated generation with human verification, we create 1,005 challenging examples mirroring real-world complexity. Evaluating 25 state-of-the-art MLLMs reveals a profound capability gap: the best model achieves only 45.1% accuracy, while most score 0% on rejection tasks. We explore two strategies for improvements: (1) test-time scaling selects optimal response by thinking trajectory to improve overall performance by up to 4.5%, and (2) data-mixture training boosts rejection accuracy from 0% to 27.9%. GroundingME thus serves as both a diagnostic tool revealing current limitations in MLLMs and a roadmap toward human-level visual grounding. Project page: https://groundingme.github.io

📄 PDF Abstract BibTeX arXiv:2512.17495

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models

2024-03-28 · Jiaxing Chen, Yuxuan Liu, Dehu Li, Xiang An 외

The rise of Multimodal Large Language Models (MLLMs), renowned for their advanced instruction-following and reasoning capabilities, has significantly propelled the field of visual reasoning. However, due to limitations i…

Instruction FollowingVisual Reasoning

Adversarial Robustness for Visual Grounding of Multimodal Large Language Models

2024-05-16 · Kuofeng Gao, Yang Bai, Jiawang Bai, Yong Yang 외

Multi-modal Large Language Models (MLLMs) have recently achieved enhanced performance across various vision-language tasks including visual grounding capabilities. However, the adversarial robustness of visual grounding …

Adversarial AttackAdversarial RobustnessReferring ExpressionReferring Expression Comprehension+1

How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical Images

2026-03-15 · Guimeng Liu, Tianze Yu, Somayeh Ebrahimkhani, Lin Zhi Zheng Shawn 외 arxiv

Generalist multimodal large language models (MLLMs) have achieved impressive performance across a wide range of vision-language tasks. However, their performance on medical tasks, particularly in zero-shot settings where…

Visual Grounding

OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models

2026-03-10 · Tengjin Weng, Wenhao Jiang, Jingyi Wang, Ming Li 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision language tasks. However, their ability in low-level visual perception, particularly in detecting fine-grained vi…

Reinforcement Learning

MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs

2024-10-16 · Yunqiu Xu, Linchao Zhu, Yi Yang

While multimodal large language models (MLLMs) have demonstrated extraordinary vision-language understanding capabilities and shown potential to serve as general-purpose assistants, their abilities to solve instance-leve…

Visual Grounding