paper-with-me

홈 › Papers

'What are you referring to?' Evaluating the Ability of Multi-Modal Dialogue Models to Process Clarificational Exchanges

2023-07-28 · Javier Chiyah-Garcia, Alessandro Suglia, Arash Eshghi, Helen Hastie

Referential ambiguities arise in dialogue when a referring expression does not uniquely identify the intended referent for the addressee. Addressees usually detect such ambiguities immediately and work with the speaker to repair it using meta-communicative, Clarificational Exchanges (CE): a Clarification Request (CR) and a response. Here, we argue that the ability to generate and respond to CRs imposes specific constraints on the architecture and objective functions of multi-modal, visually grounded dialogue models. We use the SIMMC 2.0 dataset to evaluate the ability of different state-of-the-art model architectures to process CEs, with a metric that probes the contextual updates that arise from them in the model. We find that language-based models are able to encode simple multi-modal semantic information and process some CEs, excelling with those related to the dialogue history, whilst multi-modal models can use additional learning objectives to obtain disentangled object representations, which become crucial to handle complex referential ambiguities across modalities overall.

📄 PDF Abstract BibTeX arXiv:2307.15554

Code (1)

jchiyah/what-are-you-referring-to 공식 구현

Tasks

Referring Expression

Methods 이 논문이 사용한 방법론

Repair 설명 없음

Similar Papers 제목 키워드 기반

What and Where: An Empirical Investigation of Pointing Gestures and Descriptions in Multimodal Referring Actions

2013-08-01 · WS 2013 8 · Albert Gatt, Patrizia Paggio
Text Generation

Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking

2024-12-17 · Wenjun Huang, Yang Ni, Hanning Chen, Yirui He 외

Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to localize an arbitrary number of targets based on a language expression and continuously track them in a video. This intricate task invol…

DecoderMulti-Object TrackingObject TrackingReferring Multi-Object Tracking

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

2024-06-28 · Yuxuan Zhang, Tianheng Cheng, Rui Hu, Lei Liu 외

Segment Anything Model (SAM) has attracted widespread attention for its superior interactive segmentation capabilities with visual prompts while lacking further exploration of text prompts. In this paper, we empirically …

Interactive SegmentationLanguage ModelingLanguage ModellingReferring Expression+2

On the role of effective and referring questions in GuessWhat?!

2020-07-01 · WS 2020 7 · Mauricio Mazuecos, Alberto Testoni, Raffaella Bernardi, Luciana Benotti

Task success is the standard metric used to evaluate referential visual dialogue systems. In this paper we propose two new metrics that evaluate how each question contributes to the goal. First, we measure how effective …

Less Descriptive yet Discriminative: Quantifying the Properties of Multimodal Referring Utterances via CLIP

2022-05-01 · CMCL (ACL) 2022 5 · Ece Takmaz, Sandro Pezzelle, Raquel Fernández

In this work, we use a transformer-based pre-trained multimodal model, CLIP, to shed light on the mechanisms employed by human speakers when referring to visual entities. In particular, we use CLIP to quantify the degree…

Descriptive