paper-with-me

홈 › Papers

Jointly Reinforced User Simulator and Task-oriented Dialog System with Simplified Generative Architecture

2022-10-13 · Hong Liu, Zhijian Ou, Yi Huang, Junlan Feng

Recently, there has been progress in supervised funetuning pretrained GPT-2 to build end-to-end task-oriented dialog (TOD) systems. However, online reinforcement learning of a GPT-2 based dialog system (DS), together with a end-to-end user simulator (US), has not ever been explored. Moreover, a drawback with existing GPT-2 based TOD systems is that they mostly employ the whole dialog history as input, which brings inefficiencies in memory and compute. In this paper, we first propose Simplified Generative Architectures (SGA) for DS and US respectively, both based on GPT-2 but using shortened history. Then, we successfully develop Jointly Reinforced US and DS, called SGA-JRUD. Our DS with the proposed SGA, when only supervised trained, achieves state-of-the-art performance on MultiWOZ2.1 and is more compute-efficient in both training and generation. Extensive experiments on MultiWOZ2.1 further show the superiority of SGA-JRUD in both offline and online evaluations.

📄 PDF Abstract BibTeX arXiv:2210.06706

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Jointly Reinforced User Simulator and Task-oriented Dialog System with Simplified Generative Architecture

2022-01-16 · ACL ARR January 2022 1 · Anonymous

The large pre-training language model GPT-2 has been fine-tuned in task-oriented dialog system and achieved state-of-the-art performance on many datasets. However, there's few work of reinforcement learning on these GPT-…

Language ModelingLanguage Modellingreinforcement-learningReinforcement Learning+1

A Generative User Simulator with GPT-based Architecture and Goal State Tracking for Reinforced Multi-Domain Dialog Systems

2022-10-17 · Hong Liu, Yucheng Cai, Zhijian Ou, Yi Huang 외

Building user simulators (USs) for reinforcement learning (RL) of task-oriented dialog systems (DSs) has gained more and more attention, which, however, still faces several fundamental challenges. First, it is unclear wh…

Reinforcement Learning (RL)

Iterative Policy Learning in End-to-End Trainable Task-Oriented Neural Dialog Models

2017-09-18 · Bing Liu, Ian Lane

In this paper, we present a deep reinforcement learning (RL) framework for iterative dialog policy optimization in end-to-end task-oriented dialog systems. Popular approaches in learning dialog policy with RL include let…

Deep Reinforcement LearningReinforcement LearningReinforcement Learning (RL)

Multi-Agent Task-Oriented Dialog Policy Learning with Role-Aware Reward Decomposition

2020-04-08 · ACL 2020 6 · Ryuichi Takanobu, Runze Liang, Minlie Huang

Many studies have applied reinforcement learning to train a dialog policy and show great promise these years. One common approach is to employ a user simulator to obtain a large number of simulated user experiences for r…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Reliable LLM-based User Simulator for Task-Oriented Dialogue Systems

2024-02-20 · Ivan Sekulić, Silvia Terragni, Victor Guimarães, Nghia Khau 외

In the realm of dialogue systems, user simulation techniques have emerged as a game-changer, redefining the evaluation and enhancement of task-oriented dialogue (TOD) systems. These methods are crucial for replicating re…

Data AugmentationTask-Oriented Dialogue SystemsUser Simulation