Counterfactual Off-Policy Training for Neural Dialogue Generation
Open-domain dialogue generation suffers from the data insufficiency problem due to the vast size of potential responses. In this paper, we propose to explore potential responses by counterfactual reasoning. Given an observed response, the counterfactual reasoning model automatically infers the outcome of an alternative policy that could have been taken. The resulting counterfactual response synthesized in hindsight is of higher quality than the response synthesized from scratch. Training on the counterfactual responses under the adversarial learning framework helps to explore the high-reward area of the potential response space. An empirical study on the DailyDialog dataset shows that our approach significantly outperforms the HRED model as well as the conventional adversarial learning approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
counterfactualCounterfactual ReasoningDialogue GenerationSimilar Papers 제목 키워드 기반
Counterfactual Off-Policy Training for Neural Response Generation
Open-domain dialogue generation suffers from the data insufficiency problem due to the vast size of potential responses. In this paper, we propose to explore potential responses by counterfactual reasoning. Given an obse…
counterfactualCounterfactual ReasoningDialogue GenerationResponse Generation$C^3$: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues
Video-grounded dialogue systems aim to integrate video understanding and dialogue understanding to generate responses that are relevant to both the dialogue and video context. Most existing approaches employ deep learnin…
Contrastive LearningcounterfactualDialogue UnderstandingMultimodal Reasoning+1CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers
Dialogue state trackers have made significant progress on benchmark datasets, but their generalization capability to novel and realistic scenarios beyond the held-out conversations is less understood. We propose controll…
counterfactualDialogue State TrackingMulti-domain Dialogue State TrackingCounterfactual Data Augmentation via Perspective Transition for Open-Domain Dialogues
The construction of open-domain dialogue systems requires high-quality dialogue datasets. The dialogue data admits a wide variety of responses for a given dialogue history, especially responses with different semantics. …
counterfactualCounterfactual InferenceData AugmentationCausal Discovery and Counterfactual Reasoning to Optimize Persuasive Dialogue Policies
Tailoring persuasive conversations to users leads to more effective persuasion. However, existing dialogue systems often struggle to adapt to dynamically evolving user states. This paper presents a novel method that leve…
Causal DiscoverycounterfactualCounterfactual Reasoning