Evaluating Dialogue Generation Systems via Response Selection
Existing automatic evaluation metrics for open-domain dialogue response generation systems correlate poorly with human evaluation. We focus on evaluating response generation systems via response selection. To evaluate systems properly via response selection, we propose the method to construct response selection test sets with well-chosen false candidates. Specifically, we propose to construct test sets filtering out some types of false candidates: (i) those unrelated to the ground-truth response and (ii) those acceptable as appropriate responses. Through experiments, we demonstrate that evaluating systems via response selection with the test sets developed by our method correlates more strongly with human evaluation, compared with widely used automatic evaluation metrics such as BLEU.
Code (1)
Tasks
Dialogue GenerationResponse GenerationSimilar Papers 제목 키워드 기반
A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation
Knowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support k…
Dialogue GenerationResponse GenerationPneg: Prompt-based Negative Response Generation for Dialogue Response Selection Task
In retrieval-based dialogue systems, a response selection model acts as a ranker to select the most appropriate response among several candidates. However, such selection models tend to rely on context-response content s…
Language ModelingLanguage ModellingResponse GenerationRetrievalDialogueScore: Evaluating Responses in Task-Oriented Dialogue
Task-Oriented Dialogue systems have been widely deployed in real-world applications in the last few years.Yet, evaluations of task-oriented dialogue systems are relatively limited.The informative and success score only c…
Natural Language InferenceTask-Oriented Dialogue SystemsA Compare Aggregate Transformer for Understanding Document-grounded Dialogue
Unstructured documents serving as external knowledge of the dialogues help to generate more informative responses. Previous research focused on knowledge selection (KS) in the document with dialogue. However, dialogue hi…
Response GenerationWell Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded Dialogue
Accurate knowledge selection is critical in knowledge-grounded dialogue systems. Towards a closer look at it, we offer a novel perspective to organize existing literature, i.e., knowledge selection coupled with, after, a…
Response Generation