paper-with-me

홈 › Papers

Towards Unified Dialogue System Evaluation: A Comprehensive Analysis of Current Evaluation Protocols

2020-06-10 · SIGDIAL (ACL) 2020 7 · Sarah E. Finch, Jinho D. Choi

As conversational AI-based dialogue management has increasingly become a trending topic, the need for a standardized and reliable evaluation procedure grows even more pressing. The current state of affairs suggests various evaluation protocols to assess chat-oriented dialogue management systems, rendering it difficult to conduct fair comparative studies across different approaches and gain an insightful understanding of their values. To foster this research, a more robust evaluation protocol must be set in place. This paper presents a comprehensive synthesis of both automated and human evaluation methods on dialogue systems, identifying their shortcomings while accumulating evidence towards the most effective evaluation dimensions. A total of 20 papers from the last two years are surveyed to analyze three types of evaluation protocols: automated, static, and interactive. Finally, the evaluation dimensions used in these papers are compared against our expert evaluation on the system-user dialogue data collected from the Alexa Prize 2020.

📄 PDF Abstract BibTeX arXiv:2006.06110

Code (0)

등록된 구현이 없습니다.

Tasks

Dialogue ManagementManagement

Similar Papers 제목 키워드 기반

ConvLab-3: A Flexible Dialogue System Toolkit Based on a Unified Data Format

2022-11-30 · Qi Zhu, Christian Geishauser, Hsien-Chin Lin, Carel van Niekerk 외

Task-oriented dialogue (TOD) systems function as digital assistants, guiding users through various tasks such as booking flights or finding restaurants. Existing toolkits for building TOD systems often fall short of in d…

Reinforcement Learning (RL)Transfer Learning

A Unified Pre-training Framework for Conversational AI

2021-05-06 · Siqi Bao, Bingjin Chen, Huang He, Xin Tian 외

In this work, we explore the application of PLATO-2 on various dialogue systems, including open-domain conversation, knowledge grounded dialogue, and task-oriented conversation. PLATO-2 is initially designed as an open-d…

ChatbotInteractive Evaluation of DialogResponse Generation

A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators

2023-12-24 · Chen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang 외

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggeste…

Dialogue Evaluation

LLM-Mini-CEX: Automatic Evaluation of Large Language Model for Diagnostic Conversation

2023-08-15 · Xiaoming Shi, Jie Xu, Jinru Ding, Jiali Pang 외

There is an increasing interest in developing LLMs for medical diagnosis to improve diagnosis efficiency. Despite their alluring technological potential, there is no unified and comprehensive evaluation criterion, leadin…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model+1

ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents

2025-08-28 · Tianjian Liu, Fanqi Wan, Jiajian Guo, Xiaojun Quan arxiv

Proactive dialogue has emerged as a critical and challenging research problem in advancing large language models (LLMs). Existing works predominantly focus on domain-specific or task-oriented scenarios, which leads to fr…