paper-with-me

홈 › Papers

Towards Understanding Visual Grounding in Visual Language Models

2025-09-12 · Georgios Pantazopoulos, Eda B. Özyiğit arxiv

Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide range of applications in various domains, including referring expression comprehension, answering questions pertinent to fine-grained details in images or videos, caption visual context by explicitly referring to entities, as well as low and high-level control in simulated and real environments. In this survey paper, we review representative works across the key areas of research on modern general-purpose vision language models (VLMs). We first outline the importance of grounding in VLMs, then delineate the core components of the contemporary paradigm for developing grounded models, and examine their practical applications, including benchmarks and evaluation metrics for grounded multimodal generation. We also discuss the multifaceted interrelations among visual grounding, multimodal chain-of-thought, and reasoning in VLMs. Finally, we analyse the challenges inherent to visual grounding and suggest promising directions for future research.

📄 PDF Abstract BibTeX arXiv:2509.10345

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generationReferring ExpressionVisual Grounding

Similar Papers 제목 키워드 기반

Learning to Ground VLMs without Forgetting

2024-10-14 · Aritra Bhowmik, Mohammad Mahdi Derakhshani, Dennis Koelma, Martin R. Oswald 외

Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Visual Language Models (VLMs) struggle at this task. In this paper, we introduce LynX, a framew…

DecoderLanguage ModellingMixture-of-ExpertsMultimodal Reasoning+3

Investigating Compositional Challenges in Vision-Language Models for Visual Grounding

2024-01-01 · CVPR 2024 1 · Yunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie 외

Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performan…

AttributeRelationVisual Grounding

DSM: Building A Diverse Semantic Map for 3D Visual Grounding

2025-04-11 · Qinghongbing Xie, Zijian Liang, Long Zeng

In recent years, with the growing research and application of multimodal large language models (VLMs) in robotics, there has been an increasing trend of utilizing VLMs for robotic scene understanding tasks. Existing appr…

3D visual groundingScene UnderstandingSemantic SegmentationVisual Grounding

EGM: Efficient Visual Grounding Language Models

2026-01-20 · Guanqi Zhan, Changye Li, Zhijian Liu, Yao Lu 외 arxiv

Visual grounding is an essential capability of Visual Language Models (VLMs) to understand the real physical world. Previous state-of-the-art grounding visual language models usually have large model sizes, making them h…

Visual Grounding

Like a bilingual baby: The advantage of visually grounding a bilingual language model

2022-10-11 · Khai-Nguyen Nguyen, Zixin Tang, Ankur Mali, Alex Kelly

Unlike most neural language models, humans learn language in a rich, multi-sensory and, often, multi-lingual environment. Current language models typically fail to fully capture the complexities of multilingual language …

Language ModelingLanguage ModellingSemantic SimilaritySemantic Textual Similarity+1