Reference-free Evaluation Metrics for Text Generation: A Survey
A number of automatic evaluation metrics have been proposed for natural language generation systems. The most common approach to automatic evaluation is the use of a reference-based metric that compares the model's output with gold-standard references written by humans. However, it is expensive to create such references, and for some tasks, such as response generation in dialogue, creating references is not a simple matter. Therefore, various reference-free metrics have been developed in recent years. In this survey, which intends to cover the full breadth of all NLG tasks, we investigate the most commonly used approaches, their application, and their other uses beyond evaluating models. The survey concludes by highlighting some promising directions for future research.
Code (0)
등록된 구현이 없습니다.
Tasks
Response GenerationSurveyText GenerationSimilar Papers 제목 키워드 기반
Spurious Correlations in Reference-Free Evaluation of Text Generation
Model-based, reference-free evaluation metrics have been proposed as a fast and cost-effective approach to evaluate Natural Language Generation (NLG) systems. Despite promising recent results, we find evidence that refer…
Abstractive Text SummarizationText GenerationText SummarizationOn the Limitations of Reference-Free Evaluations of Generated Text
There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to colle…
DiagnosticMachine TranslationOn the Evaluation Metrics for Paraphrase Generation
In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom: (1) Reference-free metrics achieve better performance than their reference-based counterparts…
Machine TranslationParaphrase GenerationCTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation
Existing reference-free metrics have obvious limitations for evaluating controlled text generation models. Unsupervised metrics can only provide a task-agnostic evaluation result which correlates weakly with human judgme…
Language ModelingLanguage ModellingText GenerationText InfillingLearning to Rank Visual Stories From Human Ranking Data
Visual storytelling (VIST) is a typical vision and language task that has seen extensive development in the natural language generation research domain. However, it remains unclear whether conventional automatic evaluati…
Learning-To-RankText GenerationVisual Storytelling