paper-with-me

홈 › Papers

Seeing things or seeing scenes: Investigating the capabilities of V&L models to align scene descriptions to images

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Images can be described in terms of the objects they contain, or in terms of the types of scene or place that they instantiate. In this paper we address to what extent pretrained Vision and Language models can learn to align descriptions of both types with images. We compare 3 state-of-the-art models, VisualBERT, LXMERT and CLIP. We find that (i) V\&L models are susceptible to stylistic biases acquired during pretraining; (ii) only CLIP performs consistently well on both object- and scene-level descriptions. A follow-up ablation study shows that CLIP uses object-level information in the visual modality to align with scene-level textual descriptions.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Object

Methods 이 논문이 사용한 방법론

LXMERT LXMERT is a model for learning vision-and-language cross-modality representations. It consists of a Transformer model that consists three encoders: object relationship encoder, a…
VisualBERT VisualBERT aims to reuse self-attention to implicitly align elements of the input text and regions in the input image. Visual embeddings are used to model images where the…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Are You a Racist or Am I Seeing Things? Annotator Influence on Hate Speech Detection on Twitter

2016-11-01 · WS 2016 11 · Zeerak Waseem
Hate Speech Detection

Vision: looking and seeing through our brain's information bottleneck

2025-03-24 · Li Zhaoping

Our brain recognizes only a tiny fraction of sensory input, due to an information processing bottleneck. This blinds us to most visual inputs. Since we are blind to this blindness, only a recent framework highlights this…

Seeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal

2026-02-03 · Rio Aguina-Kang, Kevin James Blackburn-Matzen, Thibault Groueix, Vladimir Kim 외 arxiv

We present SeeingThroughClutter, a method for reconstructing structured 3D representations from single images by segmenting and modeling objects individually. Prior approaches rely on intermediate tasks such as semantic …

Semantic SegmentationDepth Estimation

Seeing Faces in Things: A Model and Dataset for Pareidolia

2024-09-24 · Mark Hamilton, Simon Stent, Vasha DuTell, Anne Harrington 외

The human visual system is well-tuned to detect faces of all shapes and sizes. While this brings obvious survival advantages, such as a better chance of spotting unknown predators in the bush, it also leads to spurious f…

Single Image 3D Without a Single 3D Image

2015-12-01 · ICCV 2015 12 · David F. Fouhey, Wajahat Hussain, Abhinav Gupta, Martial Hebert

Do we really need 3D labels in order to learn how to predict 3D? In this paper, we show that one can learn a mapping from appearance to 3D properties without ever seeing a single explicit 3D label. Rather than use explic…

Scene Understanding