The use of rating and Likert scales in Natural Language Generation human evaluation tasks: A review and some recommendations
Rating and Likert scales are widely used in evaluation experiments to measure the quality of Natural Language Generation (NLG) systems. We review the use of rating and Likert scales for NLG evaluation tasks published in NLG specialized conferences over the last ten years (135 papers in total). Our analysis brings to light a number of deviations from good practice in their use. We conclude with some recommendations about the use of such scales. Our aim is to encourage the appropriate use of evaluation methodologies in the NLG community.
Code (0)
등록된 구현이 없습니다.
Tasks
nlg evaluationText GenerationSimilar Papers 제목 키워드 기반
Understanding the Impact of Experiment Design for Evaluating Dialogue System Output
Evaluation of output from natural language generation (NLG) systems is typically conducted via crowdsourced human judgments. To understand the impact of how experiment design might affect the quality and consistency of s…
Text GenerationChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
As AI becomes more integral in our lives, the need for transparency and responsibility grows. While natural language explanations (NLEs) are vital for clarifying the reasoning behind AI decisions, evaluating them through…
InformativenessUsing Natural Language Explanations to Rescale Human Judgments
The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation. A common practice is to label data via consensus an…
Question AnsweringResponse Style Characterization for Repeated Measures Using the Visual Analogue Scale
Self-report measures (e.g., Likert scales) are widely used to evaluate subjective health perceptions. Recently, the visual analog scale (VAS), a slider-based scale, has become popular owing to its ability to precisely an…
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to considerable enhancements in productivity a…
Language ModelingLanguage ModellingLarge Language ModelMultiple-choice+2