Neural User Simulation for Corpus-based Policy Optimisation of Spoken Dialogue Systems
User Simulators are one of the major tools that enable offline training of task-oriented dialogue systems. For this task the Agenda-Based User Simulator (ABUS) is often used. The ABUS is based on hand-crafted rules and its output is in semantic form. Issues arise from both properties such as limited diversity and the inability to interface a text-level belief tracker. This paper introduces the Neural User Simulator (NUS) whose behaviour is learned from a corpus and which generates natural language, hence needing a less labelled dataset than simulators generating a semantic output. In comparison to much of the past work on this topic, which evaluates user simulators on corpus-based metrics, we use the NUS to train the policy of a reinforcement learning based Spoken Dialogue System. The NUS is compared to the ABUS by evaluating the policies that were trained using the simulators. Cross-model evaluation is performed i.e. training on one simulator and testing on the other. Furthermore, the trained policies are tested on real users. In both evaluation tasks the NUS outperformed the ABUS.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialogue ManagementDiversityReinforcement LearningSpoken Dialogue SystemsTask-Oriented Dialogue SystemsUser SimulationSimilar Papers 제목 키워드 기반
Neural User Simulation for Corpus-based Policy Optimisation for Spoken Dialogue Systems
User Simulators are one of the major tools that enable offline training of task-oriented dialogue systems. For this task the Agenda-Based User Simulator (ABUS) is often used. The ABUS is based on hand-crafted rules and i…
DiversityReinforcement LearningSpoken Dialogue SystemsTask-Oriented Dialogue Systems+1Adversarial learning of neural user simulators for dialogue policy optimisation
Reinforcement learning based dialogue policies are typically trained in interaction with a user simulator. To obtain an effective and robust policy, this simulator should generate user behaviour that is both realistic an…
On-line Active Reward Learning for Policy Optimisation in Spoken Dialogue Systems
The ability to compute an accurate reward function is essential for optimising a dialogue policy via reinforcement learning. In real-world applications, using explicit user feedback as the reward signal is often unreliab…
Active LearningDecoderReinforcement LearningSpoken Dialogue SystemsThe Twins Corpus of Museum Visitor Questions
The Twins corpus is a collection of utterances spoken in interactions with two virtual characters who serve as guides at the Museum of Science in Boston. The corpus contains about 200,000 spoken utterances from museum vi…
Dialogue ManagementNatural Language Understandingspeech-recognitionSpeech Recognition+1Dialogue Strategy Adaptation to New Action Sets Using Multi-dimensional Modelling
A major bottleneck for building statistical spoken dialogue systems for new domains and applications is the need for large amounts of training data. To address this problem, we adopt the multi-dimensional approach to dia…
Dialogue ManagementManagementSpoken Dialogue SystemsTransfer Learning