Two can play this Game: Visual Dialog with Discriminative Question Generation and Answering
Human conversation is a complex mechanism with subtle nuances. It is hence an ambitious goal to develop artificial intelligence agents that can participate fluently in a conversation. While we are still far from achieving this goal, recent progress in visual question answering, image captioning, and visual question generation shows that dialog systems may be realizable in the not too distant future. To this end, a novel dataset was introduced recently and encouraging results were demonstrated, particularly for question answering. In this paper, we demonstrate a simple symmetric discriminative baseline, that can be applied to both predicting an answer as well as predicting a question. We show that this method performs on par with the state of the art, even memory net based methods. In addition, for the first time on the visual dialog dataset, we assess the performance of a system asking questions, and demonstrate how visual dialog can be generated from discriminative question generation and question answering.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningQuestion AnsweringQuestion GenerationQuestion-GenerationVisual DialogVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
GuessWhat?! Visual object discovery through multi-modal dialogue
We introduce GuessWhat?!, a two-player guessing game as a testbed for research on the interplay of computer vision and dialogue systems. The goal of the game is to locate an unknown object in a rich image scene by asking…
ObjectObject DiscoverySpatial ReasoningMulti-Modal Dialogue State Tracking for Playing GuessWhich Game
GuessWhich is an engaging visual dialogue game that involves interaction between a Questioner Bot (QBot) and an Answer Bot (ABot) in the context of image-guessing. In this game, QBot's objective is to locate a concealed …
Dialogue State TrackingVisual ReasoningLearning Better Visual Dialog Agents with Pretrained Visual-Linguistic Representation
GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in an image, based on answers from player B …
Referring ExpressionReferring Expression ComprehensionVisual DialogVisual GroundingA Visually-Aware Conversational Robot Receptionist
Socially Assistive Robots (SARs) have the potential to play an increasingly important role in a variety of contexts including healthcare, but most existing systems have very limited interactive capabilities. We will demo…
Question AnsweringDialogue Policies for Learning Board Games through Multimodal Communication
This paper presents MDP policy learning for agents to learn strategic behavior–how to play board games–during multimodal dialogues. Policies are trained offline in simulation, with dialogues carried out in a formal langu…
Board GamesInformativeness