paper-with-me

홈 › Papers

How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation

2016-03-25 · EMNLP 2016 11 · Chia-Wei Liu, Ryan Lowe, Iulian V. Serban, Michael Noseworthy, Laurent Charlin, Joelle Pineau

We investigate evaluation metrics for dialogue response generation systems where supervised labels, such as task completion, are not available. Recent works in response generation have adopted metrics from machine translation to compare a model's generated response to a single target response. We show that these metrics correlate very weakly with human judgements in the non-technical Twitter domain, and not at all in the technical Ubuntu domain. We provide quantitative and qualitative results highlighting specific weaknesses in existing metrics, and provide recommendations for future development of better automatic evaluation metrics for dialogue systems.

📄 PDF Abstract BibTeX arXiv:1603.08023

Code (2)

23LuZ/the-Evaluation-of-ChitChat-System
piekey1994/IOM pytorch

Tasks

Machine TranslationResponse GenerationTranslation

Similar Papers 제목 키워드 기반

Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue Systems

2022-12-18 · Sarah E. Finch, James D. Finch, Jinho D. Choi

Despite tremendous advancements in dialogue systems, stable evaluation still requires human judgments producing notoriously high-variance metrics due to their inherent subjectivity. Moreover, methods and labels in dialog…

ChatbotDialogue Evaluation

How to Evaluate Your Dialogue Models: A Review of Approaches

2021-08-03 · Xinmeng Li, Wansen Wu, Long Qin, Quanjun Yin

Evaluating the quality of a dialogue system is an understudied problem. The recent evolution of evaluation method motivated this survey, in which an explicit and comprehensive analysis of the existing methods is sought. …

Your Students Don't Use LLMs Like You Wish They Did

2026-04-26 · Sebastian Kobler, Matthew Clemson, Angela Sun, Jonathan K. Kummerfeld arxiv

Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals. We introduce six computational metrics for automated evaluation …

Dialogue Evaluation

Mimic and Rephrase: Reflective Listening in Open-Ended Dialogue

2019-11-01 · CONLL 2019 11 · Justin Dieter, Tian Wang, Arun Tejasvi Chaganty, Gabor Angeli 외

Reflective listening{--}demonstrating that you have heard your conversational partner{--}is key to effective communication. Expert human communicators often mimic and rephrase their conversational partner, e.g., when res…

How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics

2020-08-24 · Prasanna Parthasarathi, Joelle Pineau, Sarath Chandar

Though generative dialogue modeling is widely seen as a language modeling task, the task demands an agent to have a complex natural language understanding of its input text to carry a meaningful interaction with an user.…

Language ModelingLanguage ModellingNatural Language Understanding