Towards Understanding Visual Grounding in Visual Language Models
Visual grounding refers to the ability of a model to identify a region within some visual input that matches a textual description. Consequently, a model equipped with visual grounding capabilities can target a wide range of applications in various domains, including referring expression comprehension, answering questions pertinent to fine-grained details in images or videos, caption visual context by explicitly referring to entities, as well as low and high-level control in simulated and real environments. In this survey paper, we review representative works across the key areas of research on modern general-purpose vision language models (VLMs). We first outline the importance of grounding in VLMs, then delineate the core components of the contemporary paradigm for developing grounded models, and examine their practical applications, including benchmarks and evaluation metrics for grounded multimodal generation. We also discuss the multifaceted interrelations among visual grounding, multimodal chain-of-thought, and reasoning in VLMs. Finally, we analyse the challenges inherent to visual grounding and suggest promising directions for future research.
Code (0)
등록된 구현이 없습니다.
Tasks
multimodal generationReferring ExpressionVisual GroundingSimilar Papers 제목 키워드 기반
Learning to Ground VLMs without Forgetting
Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Visual Language Models (VLMs) struggle at this task. In this paper, we introduce LynX, a framew…
DecoderLanguage ModellingMixture-of-ExpertsMultimodal Reasoning+3Investigating Compositional Challenges in Vision-Language Models for Visual Grounding
Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performan…
AttributeRelationVisual GroundingDSM: Building A Diverse Semantic Map for 3D Visual Grounding
In recent years, with the growing research and application of multimodal large language models (VLMs) in robotics, there has been an increasing trend of utilizing VLMs for robotic scene understanding tasks. Existing appr…
3D visual groundingScene UnderstandingSemantic SegmentationVisual GroundingEGM: Efficient Visual Grounding Language Models
Visual grounding is an essential capability of Visual Language Models (VLMs) to understand the real physical world. Previous state-of-the-art grounding visual language models usually have large model sizes, making them h…
Visual GroundingLike a bilingual baby: The advantage of visually grounding a bilingual language model
Unlike most neural language models, humans learn language in a rich, multi-sensory and, often, multi-lingual environment. Current language models typically fail to fully capture the complexities of multilingual language …
Language ModelingLanguage ModellingSemantic SimilaritySemantic Textual Similarity+1