paper-with-me

홈 › Papers

Re-evaluating ADEM: A Deeper Look at Scoring Dialogue Responses

2019-02-23 · Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra, Mukundhan Srinivasan

Automatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. ADEM(Lowe et al. 2017) formulated the automatic evaluation of dialogue systems as a learning problem and showed that such a model was able to predict responses which correlate significantly with human judgements, both at utterance and system level. Their system was shown to have beaten word-overlap metrics such as BLEU with large margins. We start with the question of whether an adversary can game the ADEM model. We design a battery of targeted attacks at the neural network based ADEM evaluation system and show that automatic evaluation of dialogue systems still has a long way to go. ADEM can get confused with a variation as simple as reversing the word order in the text! We report experiments on several such adversarial scenarios that draw out counterintuitive scores on the dialogue responses. We take a systematic look at the scoring function proposed by ADEM and connect it to linear system theory to predict the shortcomings evident in the system. We also devise an attack that can fool such a system to rate a response generation system as favorable. Finally, we allude to future research directions of using the adversarial attacks to design a truly automated dialogue evaluation system.

📄 PDF Abstract BibTeX arXiv:1902.08832

Code (0)

등록된 구현이 없습니다.

Tasks

Dialogue EvaluationResponse Generation

Similar Papers 제목 키워드 기반

Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses

2017-08-23 · ACL 2017 7 · Ryan Lowe, Michael Noseworthy, Iulian V. Serban, Nicolas Angelard-Gontier 외

Automatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. Unfortunately, existing automatic evaluation metrics are biased and correlate very poorly with human judgement…

Dialogue Evaluation

RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing

2025-07-27 · Hao Xiang, Tianyi Tang, Yang Su, Bowen Yu 외 arxiv

Recent advancements in Large Language Models (LLMs) have shown outstanding potential for role-playing applications. Evaluating these capabilities is becoming crucial yet remains challenging. Existing benchmarks mostly ad…

Treating Dialogue Quality Evaluation as an Anomaly Detection Problem

2020-05-01 · LREC 2020 5 · Rostislav Nedelchev, Ricardo Usbeck, Jens Lehmann

Dialogue systems for interaction with humans have been enjoying increased popularity in the research and industry fields. To this day, the best way to estimate their success is through means of human evaluation and not a…

Anomaly DetectionDialogue Evaluation

Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective

2026-03-15 · Tianyi Zhang, David Traum arxiv

In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by coherence, consistency, and shared understanding. However, many systems fo…

Response Generation

Augmenting the Author: Exploring the Potential of AI Collaboration in Academic Writing

2024-04-23 · Joseph Tu, Hilda Hadan, Derrick M. Wang, Sabrina A Sgandurra 외

This workshop paper presents a critical examination of the integration of Generative AI (Gen AI) into the academic writing process, focusing on the use of AI as a collaborative tool. It contrasts the performance and inte…