Less Descriptive yet Discriminative: Quantifying the Properties of Multimodal Referring Utterances via CLIP
In this work, we use a transformer-based pre-trained multimodal model, CLIP, to shed light on the mechanisms employed by human speakers when referring to visual entities. In particular, we use CLIP to quantify the degree of descriptiveness (how well an utterance describes an image in isolation) and discriminativeness (to what extent an utterance is effective in picking out a single image among similar images) of human referring utterances within multimodal dialogues. Overall, our results show that utterances become less descriptive over time while their discriminativeness remains unchanged. Through analysis, we propose that this trend could be due to participants relying on the previous mentions in the dialogue history, as well as being able to distill the most discriminative information from the visual context. In general, our study opens up the possibility of using this and similar models to quantify patterns in human data and shed light on the underlying cognitive mechanisms.
Code (1)
Tasks
DescriptiveSimilar Papers 제목 키워드 기반
Zero-Shot Visual Grounding of Referring Utterances in Dialogue
This work explores whether current pretrained multimodal models, which are optimized to align images and captions, can be applied to the rather different domain of referring expressions. In particular, we test whether on…
DescriptiveVisual GroundingDiscriminative Perception via Anchored Description for Reasoning Segmentation
Reasoning segmentation increasingly employs reinforcement learning to generate explanatory reasoning chains that guide Multimodal Large Language Models. While these geometric rewards are primarily confined to guiding the…
Reinforcement LearningSoda: An Object-Oriented Functional Language for Specifying Human-Centered Problems
We present Soda (Symbolic Objective Descriptive Analysis), a language that helps to treat qualities and quantities in a natural way and greatly simplifies the task of checking their correctness. We present key properties…
DescriptiveA Tale of Three Probabilistic Families: Discriminative, Descriptive and Generative Models
The pattern theory of Grenander is a mathematical framework where patterns are represented by probability models on random variables of algebraic structures. In this paper, we review three families of probability models,…
DescriptiveFast kernel half-space depth for data with non-convex supports
Data depth is a statistical function that generalizes order and quantiles to the multivariate setting and beyond, with applications spanning over descriptive and visual statistics, anomaly detection, testing, etc. The ce…
Anomaly DetectionDescriptive