paper-with-me

홈 › Papers

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

2025-12-14 · Zhou Tao, Shida Wang, Yongxiang Hua, Haoyu Cao, Linli Xu arxiv

Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introduce DiG (Differential Grounding), a novel proxy task framework where MLLMs learn fine-grained perception by identifying and localizing all differences between similar image pairs without prior knowledge of their number. To support scalable training, we develop an automated 3D rendering-based data generation pipeline that produces high-quality paired images with fully controllable discrepancies. To address the sparsity of difference signals, we further employ curriculum learning that progressively increases complexity from single to multiple differences, enabling stable optimization. Extensive experiments demonstrate that DiG significantly improves model performance across a variety of visual perception benchmarks and that the learned fine-grained perception skills transfer effectively to standard downstream tasks, including RefCOCO, RefCOCO+, RefCOCOg, and general multimodal perception benchmarks. Our results highlight differential grounding as a scalable and robust approach for advancing fine-grained visual reasoning in MLLMs.

📄 PDF Abstract BibTeX arXiv:2512.12633

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

2025-08-23 · Leilei Guo, Antonio Carlos Rivera, Peiyu Tang, Haoxuan Ren 외 arxiv

Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) have achieved remarkable progress in natural language processing and multimodal understanding. Despite their impressive generalization capabilities, c…

Referring ExpressionVisual Reasoning

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception

2026-04-14 · Jilong Zhu, Yang Feng arxiv

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding, they frequently falter in fine-grained perception tasks that require identifying tiny objects or d…

Visual Grounding

Image Difference Grounding with Natural Language

2025-04-02 · Wenxuan Wang, Zijia Zhao, Yisi Zhang, Yepeng Tang 외

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretations. This limits their applicability in…

Visual Grounding

GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents

2025-09-19 · Xianhang Ye, Yiqing Li, Wei Dai, Miancan Liu 외 arxiv

Existing GUI grounding methods often struggle with fine-grained localization in high-resolution screenshots. To address this, we propose GUI-ARP, a novel framework that enables adaptive multi-stage inference. Equipped wi…

Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception

2025-09-02 · Changshi Zhou, Haichuan Xu, Ningquan Gu, Zhipeng Wang 외 arxiv

Language-guided long-horizon manipulation of deformable objects presents significant challenges due to high degrees of freedom, complex dynamics, and the need for accurate vision-language grounding. In this work, we focu…

Visual Grounding