paper-with-me

홈 › Papers

Less Descriptive yet Discriminative: Quantifying the Properties of Multimodal Referring Utterances via CLIP

2022-05-01 · CMCL (ACL) 2022 5 · Ece Takmaz, Sandro Pezzelle, Raquel Fernández

In this work, we use a transformer-based pre-trained multimodal model, CLIP, to shed light on the mechanisms employed by human speakers when referring to visual entities. In particular, we use CLIP to quantify the degree of descriptiveness (how well an utterance describes an image in isolation) and discriminativeness (to what extent an utterance is effective in picking out a single image among similar images) of human referring utterances within multimodal dialogues. Overall, our results show that utterances become less descriptive over time while their discriminativeness remains unchanged. Through analysis, we propose that this trend could be due to participants relying on the previous mentions in the dialogue history, as well as being able to distill the most discriminative information from the visual context. In general, our study opens up the possibility of using this and similar models to quantify patterns in human data and shed light on the underlying cognitive mechanisms.

📄 PDF Abstract BibTeX

Code (1)

ecekt/clip-desc-disc 공식 구현 pytorch

Tasks

Descriptive

Similar Papers 제목 키워드 기반

Zero-Shot Visual Grounding of Referring Utterances in Dialogue

2021-11-16 · ACL ARR November 2021 11 · Anonymous

This work explores whether current pretrained multimodal models, which are optimized to align images and captions, can be applied to the rather different domain of referring expressions. In particular, we test whether on…

DescriptiveVisual Grounding

Discriminative Perception via Anchored Description for Reasoning Segmentation

2026-03-04 · Tao Yang, Qing Zhou, Yanliang Li, Qi Wang arxiv

Reasoning segmentation increasingly employs reinforcement learning to generate explanatory reasoning chains that guide Multimodal Large Language Models. While these geometric rewards are primarily confined to guiding the…

Reinforcement Learning

Soda: An Object-Oriented Functional Language for Specifying Human-Centered Problems

2023-10-03 · Julian Alfredo Mendez

We present Soda (Symbolic Objective Descriptive Analysis), a language that helps to treat qualities and quantities in a natural way and greatly simplifies the task of checking their correctness. We present key properties…

Descriptive

A Tale of Three Probabilistic Families: Discriminative, Descriptive and Generative Models

2018-10-09 · Ying Nian Wu, Ruiqi Gao, Tian Han, Song-Chun Zhu

The pattern theory of Grenander is a mathematical framework where patterns are represented by probability models on random variables of algebraic structures. In this paper, we review three families of probability models,…

Descriptive

Fast kernel half-space depth for data with non-convex supports

2023-12-21 · Arturo Castellanos, Pavlo Mozharovskyi, Florence d'Alché-Buc, Hicham Janati

Data depth is a statistical function that generalizes order and quantiles to the multivariate setting and beyond, with applications spanning over descriptive and visual statistics, anomaly detection, testing, etc. The ce…

Anomaly DetectionDescriptive