paper-with-me

홈 › Papers

Hierarchical Reinforcement Learning for Open-Domain Dialog

2019-09-17 · Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Rosalind Picard

Open-domain dialog generation is a challenging problem; maximum likelihood training can lead to repetitive outputs, models have difficulty tracking long-term conversational goals, and training on standard movie or online datasets may lead to the generation of inappropriate, biased, or offensive text. Reinforcement Learning (RL) is a powerful framework that could potentially address these issues, for example by allowing a dialog model to optimize for reducing toxicity and repetitiveness. However, previous approaches which apply RL to open-domain dialog generation do so at the word level, making it difficult for the model to learn proper credit assignment for long-term conversational rewards. In this paper, we propose a novel approach to hierarchical reinforcement learning, VHRL, which uses policy gradients to tune the utterance-level embedding of a variational sequence model. This hierarchical approach provides greater flexibility for learning long-term, conversational rewards. We use self-play and RL to optimize for a set of human-centered conversation metrics, and show that our approach provides significant improvements -- in terms of both human evaluation and automatic metrics -- over state-of-the-art dialog models, including Transformers.

📄 PDF Abstract BibTeX arXiv:1909.07547

Code (1)

natashamjaques/neural_chat 공식 구현 pytorch

Tasks

Hierarchical Reinforcement LearningOpen-Domain Dialogreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Which Kind Is Better in Open-domain Multi-turn Dialog,Hierarchical or Non-hierarchical Models? An Empirical Study

2020-08-07 · Tian Lan, Xian-Ling Mao, Wei Wei, He-Yan Huang

Currently, open-domain generative dialog systems have attracted considerable attention in academia and industry. Despite the success of single-turn dialog generation, multi-turn dialog generation is still a big challenge…

Sub-domain Modelling for Dialogue Management with Hierarchical Reinforcement Learning

2017-06-19 · WS 2017 8 · Paweł Budzianowski, Stefan Ultes, Pei-Hao Su, Nikola Mrkšić 외

Human conversation is inherently complex, often spanning many different topics/domains. This makes policy learning for dialogue systems very challenging. Standard flat reinforcement learning methods do not provide an eff…

Dialogue ManagementHierarchical Reinforcement LearningManagementreinforcement-learning+2

A Static and Dynamic Attention Framework for Multi Turn Dialogue Generation

2024-10-28 · Wei-Nan Zhang, Yiming Cui, Kaiyan Zhang, Yifa Wang 외

Recently, research on open domain dialogue systems have attracted extensive interests of academic and industrial researchers. The goal of an open domain dialogue system is to imitate humans in conversations. Previous wor…

Dialogue Generation

MIDAS: A Dialog Act Annotation Scheme for Open Domain HumanMachine Spoken Conversations

2021-04-01 · EACL 2021 2 · Dian Yu, Zhou Yu

Dialog act prediction in open-domain conversations is an essential language comprehension task for both dialog system building and discourse analysis. Previous dialog act schemes, such as SWBD-DAMSL, are designed mainly …

Transfer Learning

Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models

2015-07-17 · Iulian V. Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville 외

We investigate the task of building open domain, conversational dialogue systems based on large dialogue corpora using generative models. Generative models produce system responses that are autonomously generated word-by…

DecoderWord Embeddings