Show Us the Way: Learning to Manage Dialog from Demonstrations
We present our submission to the End-to-End Multi-Domain Dialog Challenge Track of the Eighth Dialog System Technology Challenge. Our proposed dialog system adopts a pipeline architecture, with distinct components for Natural Language Understanding, Dialog State Tracking, Dialog Management and Natural Language Generation. At the core of our system is a reinforcement learning algorithm which uses Deep Q-learning from Demonstrations to learn a dialog policy with the help of expert examples. We find that demonstrations are essential to training an accurate dialog policy where both state and action spaces are large. Evaluation of our Dialog Management component shows that our approach is effective - beating supervised and reinforcement learning baselines.
Code (0)
등록된 구현이 없습니다.
Tasks
dialog state trackingManagementNatural Language UnderstandingQ-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Text GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Dialog Policies from Weak Demonstrations
Deep reinforcement learning is a promising approach to training a dialog manager, but current methods struggle with the large state and action spaces of multi-domain dialog systems. Building upon Deep Q-learning from Dem…
Atari GamesDeep Reinforcement LearningQ-Learningreinforcement-learning+2Teaching Arm and Head Gestures to a Humanoid Robot through Interactive Demonstration and Spoken Instruction
We describe work in progress for training a humanoid robot to produce iconic arm and head gestures as part of task-oriented dialogue interaction. This involves the development and use of a multimodal dialog manager for n…
Gesture RecognitionCausal-aware Safe Policy Improvement for Task-oriented dialogue
The recent success of reinforcement learning's (RL) in solving complex tasks is most often attributed to its capacity to explore and exploit an environment where it has been trained. Sample efficiency is usually not an i…
Dialogue ManagementManagementText Generation[CASPI] Causal-aware Safe Policy Improvement for Task-oriented Dialogue
The recent success of reinforcement learning (RL) in solving complex tasks is often attributed to its capacity to explore and exploit an environment.Sample efficiency is usually not an issue for tasks with cheap simulato…
Dialogue ManagementManagementReinforcement Learning (RL)Learning Efficient Dialogue Policy from Demonstrations through Shaping
Training a task-oriented dialogue agent with reinforcement learning is prohibitively expensive since it requires a large volume of interactions with users. Human demonstrations can be used to accelerate learning progress…
Domain Adaptation