paper-with-me

홈 › Papers

Do Vision-Language Models Understand Compound Nouns?

2024-03-30 · Sonal Kumar, Sreyan Ghosh, S Sakshi, Utkarsh Tyagi, Dinesh Manocha

Open-vocabulary vision-language models (VLMs) like CLIP, trained using contrastive loss, have emerged as a promising new paradigm for text-to-image retrieval. However, do VLMs understand compound nouns (CNs) (e.g., lab coat) as well as they understand nouns (e.g., lab)? We curate Compun, a novel benchmark with 400 unique and commonly used CNs, to evaluate the effectiveness of VLMs in interpreting CNs. The Compun benchmark challenges a VLM for text-to-image retrieval where, given a text prompt with a CN, the task is to select the correct image that shows the CN among a pair of distractor images that show the constituent nouns that make up the CN. Next, we perform an in-depth analysis to highlight CLIPs' limited understanding of certain types of CNs. Finally, we present an alternative framework that moves beyond hand-written templates for text prompts widely used by CLIP-like models. We employ a Large Language Model to generate multiple diverse captions that include the CN as an object in the scene described by the caption. Our proposed method improves CN understanding of CLIP by 8.25% on Compun. Code and benchmark are available at: https://github.com/sonalkum/Compun

📄 PDF Abstract BibTeX arXiv:2404.00419

Code (1)

sonalkum/compun 공식 구현

Tasks

Image RetrievalLanguage ModelingLanguage ModellingLarge Language ModelRetrieval

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Multiword Expressions Dataset for Indian Languages

2016-05-01 · LREC 2016 5 · Dhirendra Singh, Sudha Bhingardive, Pushpak Bhattacharyya

Multiword Expressions (MWEs) are used frequently in natural languages, but understanding the diversity in MWEs is one of the open problem in the area of Natural Language Processing. In the context of Indian languages, MW…

DiversityPOSTAGvalid

Machine Reading with Background Knowledge

2016-12-16 · Ndapandula Nakashole, Tom M. Mitchell

Intelligent systems capable of automatically understanding natural language text are important for many artificial intelligence applications including mobile phone voice assistants, computer vision, and robotics. Underst…

Prepositional Phrase AttachmentReading Comprehension

Detection of Compound Nouns and Light Verb Constructions using IndoWordNet

2016-01-01 · GWC 2016 1 · Dhirendra Singh, Sudha Bhingardive, Pushpak Bhattacharyyaa

Detection of MultiWord Expressions (MWEs) is one of the fundamental problems in Natural Language Processing. In this paper, we focus on two categories of MWEs - Compound Nouns and Light Verb Constructions. These two cate…

Using the Web as an Implicit Training Set: Application to Noun Compound Syntax and Semantics

2019-11-23 · Preslav Nakov

An important characteristic of English written text is the abundance of noun compounds - sequences of nouns acting as a single noun, e.g., colon cancer tumor suppressor protein. While eventually mastered by domain expert…

Information RetrievalMachine TranslationPrepositional Phrase AttachmentQuestion Answering+2

Analyzing and Aligning German compound nouns

2012-05-01 · LREC 2012 5 · Marion Weller, Ulrich Heid

In this paper, we present and evaluate an approach for the compositional alignment of compound nouns using comparable corpora from technical domains. The task of term alignment consists in relating a source language term…

LemmatizationTranslation