paper-with-me

홈 › Papers

Can Feedback Enhance Semantic Grounding in Large Vision-Language Models?

2024-04-09 · Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David Acuna

Enhancing semantic grounding abilities in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying the training recipes. In this work, we venture into an orthogonal direction and explore whether VLMs can improve their semantic grounding by "receiving" feedback, without requiring in-domain data, fine-tuning, or modifications to the network architectures. We systematically analyze this hypothesis using a feedback mechanism composed of a binary signal. We find that if prompted appropriately, VLMs can utilize feedback both in a single step and iteratively, showcasing the potential of feedback as an alternative technique to improve grounding in internet-scale VLMs. Furthermore, VLMs, like LLMs, struggle to self-correct errors out-of-the-box. However, we find that this issue can be mitigated via a binary verification mechanism. Finally, we explore the potential and limitations of amalgamating these findings and applying them iteratively to automatically enhance VLMs' grounding performance, showing grounding accuracy consistently improves using automated feedback across all models in all settings investigated. Overall, our iterative framework improves semantic grounding in VLMs by more than 15 accuracy points under noise-free feedback and up to 5 accuracy points under a simple automated binary verification mechanism. The project website is hosted at https://andrewliao11.github.io/vlms_feedback

📄 PDF Abstract BibTeX arXiv:2404.06510

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?

2025-01-01 · CVPR 2025 1 · Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David Acuna

Improving semantic grounding in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying the training recipes. In this work, we venture i…

Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models

2023-09-07 · Jiaying Lu, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo 외

Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, particularly regarding safety, robustness, …

Question AnsweringVisual Question Answering

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

2026-05-15 · Yuyuan Liu, Yiping Ji, Anjie Le, Jiayuan Zhu 외 arxiv

Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grounding. However, existing methods, mainly based on GRPO, assign rewar…

Reinforcement Learning

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding

2025-05-08 · Henry Zheng, Hao Shi, Qihang Peng, Yong Xien Chng 외

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D vi…

3D visual groundingcross-modal alignmentVisual Grounding

Semantic-Spatial Discriminability Enhancement for Generalized Visual Grounding

2026-08-31 · Kaiyan Lei, Xu-Yao Zhang arxiv

Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target scenarios. Previous …

Visual Grounding