Region under Discussion for visual dialog
Visual Dialog is assumed to require the dialog history to generate correct responses during a dialog. However, it is not clear from previous work how dialog history is needed for visual dialog. In this paper we define what it means for a visual question to require dialog history and we release a subset of the Guesswhat?! questions for which their dialog history completely changes their responses. We propose a novel interpretable representation that visually grounds dialog history: the Region under Discussion. It constrains the image’s spatial features according to a semantic representation of the history inspired by the information structure notion of Question under Discussion.We evaluate the architecture on task-specific multimodal models and the visual transformer model LXMERT.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual DialogMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ICCV23 Visual-Dialog Emotion Explanation Challenge: SEU_309 Team Technical Report
The Visual-Dialog Based Emotion Explanation Generation Challenge focuses on generating emotion explanations through visual-dialog interactions in art discussions. Our approach combines state-of-the-art multi-modal models…
Explanation GenerationLanguage ModelingLanguage ModellingVisual DialogCan AI agents understand spoken conversations about data visualizations in online meetings?
In this short paper, we present work evaluating an AI agent's understanding of spoken conversations about data visualizations in an online meeting scenario. There is growing interest in the development of AI-assistants t…
MoralDial: A Framework to Train and Evaluate Moral Dialogue Systems via Moral Discussions
Morality in dialogue systems has raised great attention in research recently. A moral dialogue system aligned with users' values could enhance conversation engagement and user connections. In this paper, we propose a fra…
Annotating anaphoric phenomena in situated dialogue
In recent years several corpora have been developed for vision and language tasks. With this paper, we intend to start a discussion on the annotation of referential phenomena in situated dialogue. We argue that there is …
coreference-resolutionCoreference ResolutionDualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any…
feature selectionQuestion AnsweringVisual DialogVisual Question Answering+1