VAQUUM: Are Vague Quantifiers Grounded in Visual Data?
Vague quantifiers such as "a few" and "many" are influenced by many contextual factors, including how many objects are present in a given context. In this work, we evaluate the extent to which vision-and-language models (VLMs) are compatible with humans when producing or judging the appropriateness of vague quantifiers in visual contexts. We release a novel dataset, VAQUUM, containing 20300 human ratings on quantified statements across a total of 1089 images. Using this dataset, we compare human judgments and VLM predictions using three different evaluation methods. Our findings show that VLMs, like humans, are influenced by object counts in vague quantifier use. However, we find significant inconsistencies across models in different evaluation settings, suggesting that judging and producing vague quantifiers rely on two different processes.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Are LLMs Models of Distributional Semantics? A Case Study on Quantifiers
Distributional semantics is the linguistic theory that a word's meaning can be derived from its distribution in natural language (i.e., its use). Language models are commonly viewed as an implementation of distributional…
Linguists Who Use Probabilistic Models Love Them: Quantification in Functional Distributional Semantics
Functional Distributional Semantics provides a computationally tractable framework for learning truth-conditional semantics from a corpus. Previous work in this framework has provided a probabilistic version of first-ord…
Bayesian InferenceComparatives, Quantifiers, Proportions: A Multi-Task Model for the Learning of Quantities from Vision
The present work investigates whether different quantification mechanisms (set comparison, vague quantification, and proportional estimation) can be jointly learned from visual scenes by a multi-task computational model.…
A Compositional Bayesian Semantics for Natural Language
We propose a compositional Bayesian semantics that interprets declarative sentences in a natural language by assigning them probability conditions. These are conditional probabilities that estimate the likelihood that a …
SentenceBe Precise or Fuzzy: Learning the Meaning of Cardinals and Quantifiers from Vision
People can refer to quantities in a visual scene by using either exact cardinals (e.g. one, two, three) or natural language quantifiers (e.g. few, most, all). In humans, these two processes underlie fairly different cogn…