Knowledge Supports Visual Language Grounding: A Case Study on Colour Terms
In human cognition, world knowledge supports the perception of object colours: knowing that trees are typically green helps to perceive their colour in certain contexts. We go beyond previous studies on colour terms using isolated colour swatches and study visual grounding of colour terms in realistic objects. Our models integrate processing of visual information and object-specific knowledge via hard-coded (late) or learned (early) fusion. We find that both models consistently outperform a bottom-up baseline that predicts colour terms solely from visual inputs, but show interesting differences when predicting atypical colours of so-called colour diagnostic objects. Our models also achieve promising results when tested on new object categories not seen during training.
Code (0)
등록된 구현이 없습니다.
Tasks
DiagnosticObjectVisual GroundingWorld KnowledgeSimilar Papers 제목 키워드 기반
S-Chain: Structured Visual Chain-of-Thought For Medicine
Faithful reasoning in medical vision-language models (VLMs) requires not only accurate predictions but also transparent alignment between textual rationales and visual evidence. While Chain-of-Thought (CoT) prompting has…
Visual Question AnsweringVisual GroundingVividMed: Vision Language Model with Versatile Visual Grounding for Medicine
Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable promise in generating visually grounded responses. However, their application in the medical domain is hindered by unique challenges. For …
Language ModelingLanguage ModellingQuestion AnsweringSemantic Segmentation+3Visual Grounding of Inter-lingual Word-Embeddings
Visual grounding of Language aims at enriching textual representations of language with multiple sources of visual knowledge such as images and videos. Although visual grounding is an area of intense research, inter-ling…
Visual GroundingWord EmbeddingsWord SimilarityHiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
Visual grounding, which aims to ground a visual region via natural language, is a task that heavily relies on cross-modal alignment. Existing works utilized uni-modal pre-trained models to transfer visual or linguistic k…
cross-modal alignmentVisual GroundingGuiding Visual Question Answering with Attention Priors
The current success of modern visual reasoning systems is arguably attributed to cross-modality attention mechanisms. However, in deliberative reasoning such as in VQA, attention is unconstrained at each step, and thus m…
Question AnsweringVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)+1