paper-with-me

홈 › Papers

The use of rating and Likert scales in Natural Language Generation human evaluation tasks: A review and some recommendations

2019-10-01 · WS 2019 10 · Jacopo Amidei, Paul Piwek, Alistair Willis

Rating and Likert scales are widely used in evaluation experiments to measure the quality of Natural Language Generation (NLG) systems. We review the use of rating and Likert scales for NLG evaluation tasks published in NLG specialized conferences over the last ten years (135 papers in total). Our analysis brings to light a number of deviations from good practice in their use. We conclude with some recommendations about the use of such scales. Our aim is to encourage the appropriate use of evaluation methodologies in the NLG community.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

nlg evaluationText Generation

Similar Papers 제목 키워드 기반

Understanding the Impact of Experiment Design for Evaluating Dialogue System Output

2020-07-01 · WS 2020 7 · Sashank Santhanam, Samira Shaikh

Evaluation of output from natural language generation (NLG) systems is typically conducted via crowdsourced human judgments. To understand the impact of how experiment design might affect the quality and consistency of s…

Text Generation

ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?

2024-03-26 · Fan Huang, Haewoon Kwak, Kunwoo Park, Jisun An

As AI becomes more integral in our lives, the need for transparency and responsibility grows. While natural language explanations (NLEs) are vital for clarifying the reasoning behind AI decisions, evaluating them through…

Informativeness

Using Natural Language Explanations to Rescale Human Judgments

2023-05-24 · Manya Wadhwa, Jifan Chen, Junyi Jessy Li, Greg Durrett

The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation. A common practice is to label data via consensus an…

Question Answering

Response Style Characterization for Repeated Measures Using the Visual Analogue Scale

2024-03-15 · Shunsuke Minusa, Tadayuki Matsumura, Kanako Esaki, Yang Shao 외

Self-report measures (e.g., Likert scales) are widely used to evaluate subjective health perceptions. Recently, the visual analog scale (VAS), a slider-based scale, has become popular owing to its ability to precisely an…

SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text

2024-11-25 · Reshmi Ghosh, Tianyi Yao, Lizzy Chen, Sadid Hasan 외

Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to considerable enhancements in productivity a…

Language ModelingLanguage ModellingLarge Language ModelMultiple-choice+2