paper-with-me

홈 › Papers

RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems

2017-01-11 · Chongyang Tao, Lili Mou, Dongyan Zhao, Rui Yan

Open-domain human-computer conversation has been attracting increasing attention over the past few years. However, there does not exist a standard automatic evaluation metric for open-domain dialog systems; researchers usually resort to human annotation for model evaluation, which is time- and labor-intensive. In this paper, we propose RUBER, a Referenced metric and Unreferenced metric Blended Evaluation Routine, which evaluates a reply by taking into consideration both a groundtruth reply and a query (previous user-issued utterance). Our metric is learnable, but its training does not require labels of human satisfaction. Hence, RUBER is flexible and extensible to different datasets and languages. Experiments on both retrieval and generative dialog systems show that RUBER has a high correlation with human annotation.

📄 PDF Abstract BibTeX arXiv:1701.03079

Code (1)

thu-coai/OpenMEVA tf

Tasks

Dialogue EvaluationOpen-Domain DialogRetrieval

Similar Papers 제목 키워드 기반

Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings

2019-04-24 · WS 2019 6 · Sarik Ghazarian, Johnny Tian-Zheng Wei, Aram Galstyan, Nanyun Peng

Despite advances in open-domain dialogue systems, automatic evaluation of such systems is still a challenging problem. Traditional reference-based metrics such as BLEU are ineffective because there could be many valid re…

Dialogue EvaluationvalidWord Embeddings

uBLEU: Uncertainty-Aware Automatic Evaluation Method for Open-Domain Dialogue Systems

2020-07-01 · ACL 2020 6 · Tsuta Yuma, Naoki Yoshinaga, Masashi Toyoda

Because open-domain dialogues allow diverse responses, basic reference-based metrics such as BLEU do not work well unless we prepare a massive reference set of high-quality responses for input utterances. To reduce this …

Generating Negative Samples by Manipulating Golden Responses for Unsupervised Learning of a Response Evaluation Model

2021-06-01 · NAACL 2021 4 · ChaeHun Park, Eugene Jang, Wonsuk Yang, Jong Park

Evaluating the quality of responses generated by open-domain conversation systems is a challenging task. This is partly because there can be multiple appropriate responses to a given dialogue history. Reference-based met…

Dialogue Evaluation

Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues

2024-07-20 · Dmitrii Galimzianov, Viacheslav Vyshegorodtsev

Interruption in a dialogue occurs when the listener begins their speech before the current speaker finishes speaking. Interruptions can be broadly divided into two groups: cooperative (when the listener wants to support …

Speech Interruption Detection

USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation

2020-05-01 · ACL 2020 6 · Shikib Mehri, Maxine Eskenazi

The lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research. Standard language generation metrics have been shown to be ineffective for evaluating dialog models. To this end, th…

Dialogue EvaluationOpen-Domain DialogText Generation