Towards Best Experiment Design for Evaluating Dialogue System Output
To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human judgments can suffer from the inconsistency of ratings, extant research has also found that the design of the evaluation task affects the consistency and quality of human judgments. We conduct a between-subjects study to understand the impact of four experiment conditions on human ratings of dialogue system output. In addition to discrete and continuous scale ratings, we also experiment with a novel application of Best-Worst scaling to dialogue evaluation. Through our systematic study with 40 crowdsourced workers in each task, we find that using continuous scales achieves more consistent ratings than Likert scale or ranking-based experiment design. Additionally, we find that factors such as time taken to complete the task and no prior experience of participating in similar studies of rating dialogue system output positively impact consistency and agreement amongst raters
Code (1)
Tasks
Dialogue EvaluationSimilar Papers 제목 키워드 기반
Evaluating N-best Calibration of Natural Language Understanding for Dialogue Systems
A Natural Language Understanding (NLU) component can be used in a dialogue system to perform intent classification, returning an N-best list of hypotheses with corresponding confidence estimates. We perform an in-depth e…
Classificationintent-classificationIntent ClassificationNatural Language UnderstandingRe-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses
Automatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. ADEM(Lowe et al. 2017) formulated the automatic evaluation of dialogue systems as a learning problem and showe…
Dialogue EvaluationResponse GenerationHow To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics
Though generative dialogue modeling is widely seen as a language modeling task, the task demands an agent to have a complex natural language understanding of its input text to carry a meaningful interaction with an user.…
Language ModelingLanguage ModellingNatural Language Understanding"How Robust r u?": Evaluating Task-Oriented Dialogue Systems on Spoken Conversations
Most prior work in dialogue modeling has been on written conversations mostly because of existing data sets. However, written dialogues are not sufficient to fully capture the nature of spoken conversations as well as th…
BenchmarkingDialogue State TrackingMulti-domain Dialogue State Trackingspeech-recognition+3Treating Dialogue Quality Evaluation as an Anomaly Detection Problem
Dialogue systems for interaction with humans have been enjoying increased popularity in the research and industry fields. To this day, the best way to estimate their success is through means of human evaluation and not a…
Anomaly DetectionDialogue Evaluation