paper-with-me

Papers

Towards Best Experiment Design for Evaluating Dialogue System Output

2019-09-23 · WS 2019 10 · Sashank Santhanam, Samira Shaikh

To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human judgments can suffer from the inconsistency of ratings, extant research has also found that the design of the evaluation task affects the consistency and quality of human judgments. We conduct a between-subjects study to understand the impact of four experiment conditions on human ratings of dialogue system output. In addition to discrete and continuous scale ratings, we also experiment with a novel application of Best-Worst scaling to dialogue evaluation. Through our systematic study with 40 crowdsourced workers in each task, we find that using continuous scales achieves more consistent ratings than Likert scale or ranking-based experiment design. Additionally, we find that factors such as time taken to complete the task and no prior experience of participating in similar studies of rating dialogue system output positively impact consistency and agreement amongst raters

📄 PDF Abstract BibTeX arXiv:1909.10122

Code (1)

sashank06/INLG_eval 공식 구현

Tasks

Dialogue Evaluation

Similar Papers 제목 키워드 기반

Evaluating N-best Calibration of Natural Language Understanding for Dialogue Systems

2022-09-01 · SIGDIAL (ACL) 2022 9 · Ranim Khojah, Alexander Berman, Staffan Larsson

A Natural Language Understanding (NLU) component can be used in a dialogue system to perform intent classification, returning an N-best list of hypotheses with corresponding confidence estimates. We perform an in-depth e…

Classificationintent-classificationIntent ClassificationNatural Language Understanding

Re-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses

2019-02-23 · Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra, Mukundhan Srinivasan

Automatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. ADEM(Lowe et al. 2017) formulated the automatic evaluation of dialogue systems as a learning problem and showe…

Dialogue EvaluationResponse Generation

How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics

2020-08-24 · Prasanna Parthasarathi, Joelle Pineau, Sarath Chandar

Though generative dialogue modeling is widely seen as a language modeling task, the task demands an agent to have a complex natural language understanding of its input text to carry a meaningful interaction with an user.…

Language ModelingLanguage ModellingNatural Language Understanding

"How Robust r u?": Evaluating Task-Oriented Dialogue Systems on Spoken Conversations

2021-09-28 · Seokhwan Kim, Yang Liu, Di Jin, Alexandros Papangelis 외

Most prior work in dialogue modeling has been on written conversations mostly because of existing data sets. However, written dialogues are not sufficient to fully capture the nature of spoken conversations as well as th…

BenchmarkingDialogue State TrackingMulti-domain Dialogue State Trackingspeech-recognition+3

Treating Dialogue Quality Evaluation as an Anomaly Detection Problem

2020-05-01 · LREC 2020 5 · Rostislav Nedelchev, Ricardo Usbeck, Jens Lehmann

Dialogue systems for interaction with humans have been enjoying increased popularity in the research and industry fields. To this day, the best way to estimate their success is through means of human evaluation and not a…

Anomaly DetectionDialogue Evaluation