USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation
The lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research. Standard language generation metrics have been shown to be ineffective for evaluating dialog models. To this end, this paper presents USR, an UnSupervised and Reference-free evaluation metric for dialog. USR is a reference-free metric that trains unsupervised models to measure several desirable qualities of dialog. USR is shown to strongly correlate with human judgment on both Topical-Chat (turn-level: 0.42, system-level: 1.0) and PersonaChat (turn-level: 0.48 and system-level: 1.0). USR additionally produces interpretable measures for several desirable properties of dialog.
Code (1)
Tasks
Dialogue EvaluationOpen-Domain DialogText GenerationSimilar Papers 제목 키워드 기반
Combining Structured and Unstructured Knowledge in an Interactive Search Dialogue System
Users of interactive search dialogue systems specify their preferences with natural language utterances. However, a schema-driven system is limited to handling the preferences that correspond to the predefined database c…
Semantic SimilaritySemantic Textual SimilaritySpurious Correlations in Reference-Free Evaluation of Text Generation
Model-based, reference-free evaluation metrics have been proposed as a fast and cost-effective approach to evaluate Natural Language Generation (NLG) systems. Despite promising recent results, we find evidence that refer…
Abstractive Text SummarizationText GenerationText SummarizationMeasuring the Robustness of Reference-Free Dialogue Evaluation Systems
Advancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative responses. We present a benchmark for evaluatin…
Dialogue EvaluationTAGRUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems
Open-domain human-computer conversation has been attracting increasing attention over the past few years. However, there does not exist a standard automatic evaluation metric for open-domain dialog systems; researchers u…
Dialogue EvaluationOpen-Domain DialogRetrievalReference-free Evaluation Metrics for Text Generation: A Survey
A number of automatic evaluation metrics have been proposed for natural language generation systems. The most common approach to automatic evaluation is the use of a reference-based metric that compares the model's outpu…
Response GenerationSurveyText Generation