paper-with-me

홈 › Papers

Seeing past words: Testing the cross-modal capabilities of pretrained V&L models on counting tasks

2020-12-22 · ACL (mmsr, IWCS) 2021 6 · Letitia Parcalabescu, Albert Gatt, Anette Frank, Iacer Calixto

We investigate the reasoning ability of pretrained vision and language (V&L) models in two tasks that require multimodal integration: (1) discriminating a correct image-sentence pair from an incorrect one, and (2) counting entities in an image. We evaluate three pretrained V&L models on these tasks: ViLBERT, ViLBERT 12-in-1 and LXMERT, in zero-shot and finetuned settings. Our results show that models solve task (1) very well, as expected, since all models are pretrained on task (1). However, none of the pretrained V&L models is able to adequately solve task (2), our counting probe, and they cannot generalise to out-of-distribution quantities. We propose a number of explanations for these findings: LXMERT (and to some extent ViLBERT 12-in-1) show some evidence of catastrophic forgetting on task (1). Concerning our results on the counting probe, we find evidence that all models are impacted by dataset bias, and also fail to individuate entities in the visual input. While a selling point of pretrained V&L models is their ability to solve complex tasks, our findings suggest that understanding their reasoning and grounding capabilities requires more targeted investigations on specific phenomena.

📄 PDF Abstract BibTeX arXiv:2012.12352

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceTask 2

Methods 이 논문이 사용한 방법론

LXMERT LXMERT is a model for learning vision-and-language cross-modality representations. It consists of a Transformer model that consists three encoders: object relationship encoder, a…
ViLBERT Vision-and-Language BERT (ViLBERT) is a BERT-based model for learning task-agnostic joint representations of image content and…

Similar Papers 제목 키워드 기반

Seeing Voices and Hearing Faces: Cross-modal biometric matching

2018-04-01 · CVPR 2018 6 · Arsha Nagrani, Samuel Albanie, Andrew Zisserman

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at ans…

Face RecognitionSpeaker Identification

Seeing without Pixels: Perception from Camera Trajectories

2025-11-26 · Zihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman 외 arxiv

Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. T…

Camera Pose EstimationContrastive Learning

Are words equally surprising in audio and audio-visual comprehension?

2023-07-14 · Pranava Madhyastha, Ye Zhang, Gabriella Vigliocco

We report a controlled study investigating the effect of visual information (i.e., seeing the speaker) on spoken language comprehension. We compare the ERP signature (N400) associated with each word in audio-only and aud…

ERP

Mapping the Structure and Evolution of Software Testing Research Over the Past Three Decades

2021-09-09 · Alireza Salahirad, Gregory Gay, Ehsan Mohammadi

Background: The field of software testing is growing and rapidly-evolving. Aims: Based on keywords assigned to publications, we seek to identify predominant research topics and understand how they are connected and have …

BIG-bench Machine LearningProgram Repairsoftware testing

Seeing The Words: Evaluating AI-generated Biblical Art

2025-04-23 · Hidde Makimei, Shuai Wang, Willem van Peursen

The past years witnessed a significant amount of Artificial Intelligence (AI) tools that can generate images from texts. This triggers the discussion of whether AI can generate accurate images using text from the Bible w…