paper-with-me

Papers

Towards a Metric for Automated Conversational Dialogue System Evaluation and Improvement

2019-09-26 · WS 2019 10 · Jan Deriu, Mark Cieliebak

We present "AutoJudge", an automated evaluation method for conversational dialogue systems. The method works by first generating dialogues based on self-talk, i.e. dialogue systems talking to itself. Then, it uses human ratings on these dialogues to train an automated judgement model. Our experiments show that AutoJudge correlates well with the human ratings and can be used to automatically evaluate dialogue systems, even in deployed systems. In a second part, we attempt to apply AutoJudge to improve existing systems. This works well for re-ranking a set of candidate utterances. However, our experiments show that AutoJudge cannot be applied as reward for reinforcement learning, although the metric can distinguish good from bad dialogues. We discuss potential reasons, but state here already that this is still an open question for further research.

📄 PDF Abstract BibTeX arXiv:1909.12066

Code (0)

등록된 구현이 없습니다.

Tasks

Open-Ended Question AnsweringReinforcement LearningReinforcement Learning (RL)Re-Ranking

Similar Papers 제목 키워드 기반

Measuring Conversational Fluidity in Automated Dialogue Agents

2019-10-25 · Keith Vella, Massimo Poesio, Michael Sigamani, Cihan Dogan 외

We present an automated evaluation method to measure fluidity in conversational dialogue systems. The method combines various state of the art Natural Language tools into a classifier, and human ratings on these dialogue…

Evaluating Conversational Recommender Systems with Large Language Models: A User-Centric Evaluation Framework

2025-01-16 · Nuo Chen, Quanyu Dai, Xiaoyu Dong, Xiao-Ming Wu 외

Conversational recommender systems (CRS) involve both recommendation and dialogue tasks, which makes their evaluation a unique challenge. Although past research has analyzed various factors that may affect user satisfact…

Recommendation Systems

Your Students Don't Use LLMs Like You Wish They Did

2026-04-26 · Sebastian Kobler, Matthew Clemson, Angela Sun, Jonathan K. Kummerfeld arxiv

Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals. We introduce six computational metrics for automated evaluation …

Dialogue Evaluation

Towards Automatic Evaluation of Task-Oriented Dialogue Flows

2024-11-15 · Mehrnoosh Mirtaheri, Nikhil Varghese, Chandra Khatri, Amol Kelkar

Task-oriented dialogue systems rely on predefined conversation schemes (dialogue flows) often represented as directed acyclic graphs. These flows can be manually designed or automatically generated from previously record…

Task-Oriented Dialogue Systems

LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems

2023-11-15 · Nalin Kumar, Ondřej Dušek

Linguistic entrainment, or alignment, represents a phenomenon where linguistic patterns employed by conversational participants converge to one another. While entrainment has been shown to produce a more natural user exp…

Task-Oriented Dialogue Systems