RMM: A Recursive Mental Model for Dialog Navigation
Language-guided robots must be able to both ask humans questions and understand answers. Much existing work focuses only on the latter. In this paper, we go beyond instruction following and introduce a two-agent task where one agent navigates and asks questions that a second, guiding agent answers. Inspired by theory of mind, we propose the Recursive Mental Model (RMM). The navigating agent models the guiding agent to simulate answers given candidate generated questions. The guiding agent in turn models the navigating agent to simulate navigation steps it would take to generate answers. We use the progress agents make towards the goal as a reinforcement learning reward signal to directly inform not only navigation actions, but also both question and answer generation. We demonstrate that RMM enables better generalization to novel environments. Interlocutor modelling may be a way forward for human-agent dialogue where robots need to both ask and answer questions.
Code (1)
Tasks
Answer GenerationInstruction FollowingmodelSimilar Papers 제목 키워드 기반
RMM: A Recursive Mental Model for Dialogue Navigation
Language-guided robots must be able to both ask humans questions and understand answers. Much existing work focuses only on the latter. In this paper, we go beyond instruction following and introduce a two-agent task whe…
Answer GenerationInstruction FollowingmodelRecursive Visual Attention in Visual Dialog
Visual dialog is a challenging vision-language task, which requires the agent to answer multi-round questions about an image. It typically needs to address two major problems: (1) How to answer visually-grounded question…
Question AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)Utterance Intent Classification of a Spoken Dialogue System with Efficiently Untied Recursive Autoencoders
Recursive autoencoders (RAEs) for compositionality of a vector space model were applied to utterance intent classification of a smartphone-based Japanese-language spoken dialogue system. Though the RAEs express a nonline…
Automatic Speech Recognition (ASR)ClassificationGeneral Classificationintent-classification+4DialNav: Multi-turn Dialog Navigation with a Remote Guide
We introduce DialNav, a novel collaborative embodied dialog task, where a navigation agent (Navigator) and a remote guide (Guide) engage in multi-turn dialog to reach a goal location. Unlike prior work, DialNav aims for …
Recursive Template-based Frame Generation for Task Oriented Dialog
The Natural Language Understanding (NLU) component in task oriented dialog systems processes a user{'}s request and converts it into structured information that can be consumed by downstream components such as the Dialog…
DecoderNatural Language Understanding