Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?
Improving semantic grounding in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying the training recipes. In this work, we venture into an orthogonal direction and explore self-correction in VLMs focusing on semantic grounding. We find that VLMs can correct their own semantic grounding mistakes when properly prompted and framed for the task, without any fine-tuning or even access to oracle feedback. We also introduce a self-correction framework in an iterative setting which consistently improves performance across all models investigated. Overall, we show that iterative self-correction consistently improves VLM performance in semantic grounding by up to 8.4 accuracy points across all models investigated, without requiring fine-tuning, additional architectural changes, or external data. Our exploration of self-correction also reveals that, even after several rounds of feedback, strong models like GPT-4V and GPT-4o retain limited capability in leveraging oracle feedback, suggesting promising directions for further research.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Can Feedback Enhance Semantic Grounding in Large Vision-Language Models?
Enhancing semantic grounding abilities in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying the training recipes. In this work, we ven…
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models
Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should guide robot action prediction. Yet robot fine-tuning is optimized as i…
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, particularly regarding safety, robustness, …
Question AnsweringVisual Question AnsweringProbing Semantic Grounding in Language Models of Code with Representational Similarity Analysis
Representational Similarity Analysis is a method from cognitive neuroscience, which helps in comparing representations from two different sources of data. In this paper, we propose using Representational Similarity Analy…
Numerically Grounded Language Models for Semantic Error Correction
Semantic error detection and correction is an important task for applications such as fact checking, speech-to-text or grammatical error correction. Current approaches generally focus on relatively shallow semantics and …
Fact CheckingGrammatical Error CorrectionLanguage ModelingLanguage Modelling+1