paper-with-me

Papers

ChatMatch: Evaluating Chatbots by Autonomous Chat Tournaments

2022-05-01 · ACL 2022 5 · Ruolan Yang, Zitong Li, Haifeng Tang, Kenny Zhu

Existing automatic evaluation systems of chatbots mostly rely on static chat scripts as ground truth, which is hard to obtain, and requires access to the models of the bots as a form of “white-box testing”. Interactive evaluation mitigates this problem but requires human involvement. In our work, we propose an interactive chatbot evaluation framework in which chatbots compete with each other like in a sports tournament, using flexible scoring metrics. This framework can efficiently rank chatbots independently from their model architectures and the domains for which they are trained.

📄 PDF Abstract BibTeX

Code (1)

ruolanyang/chatmatch 공식 구현 pytorch

Tasks

Chatbot

Similar Papers 제목 키워드 기반

Benchmarking LLM powered Chatbots: Methods and Metrics

2023-08-08 · Debarag Banerjee, Pooja Singh, Arjun Avadhanam, Saksham Srivastava

Autonomous conversational agents, i.e. chatbots, are becoming an increasingly common mechanism for enterprises to provide support to customers and partners. In order to rate chatbots, especially ones powered by Generativ…

BenchmarkingChatbot

LLM-empowered Chatbots for Psychiatrist and Patient Simulation: Application and Evaluation

2023-05-23 · Siyuan Chen, Mengyue Wu, Kenny Q. Zhu, Kunyao Lan 외

Empowering chatbots in the field of mental health is receiving increasing amount of attention, while there still lacks exploration in developing and evaluating chatbots in psychiatric outpatient scenarios. In this work, …

ChatbotDiagnostic

If I Hear You Correctly: Building and Evaluating Interview Chatbots with Active Listening Skills

2020-02-05 · Ziang Xiao, Michelle X. Zhou, Wenxi Chen, Huahai Yang 외

Interview chatbots engage users in a text-based conversation to draw out their views and opinions. It is, however, challenging to build effective interview chatbots that can handle user free-text responses to open-ended …

Chatbot

Evaluating Chatbots to Promote Users' Trust -- Practices and Open Problems

2023-09-09 · Biplav Srivastava, Kausik Lakkaraju, Tarmo Koppel, Vignesh Narayanan 외

Chatbots, the common moniker for collaborative assistants, are Artificial Intelligence (AI) software that enables people to naturally interact with them to get tasks done. Although chatbots have been studied since the da…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model

Evaluator for Emotionally Consistent Chatbots

2021-12-02 · Chenxiao Liu, Guanzhi Deng, Tao Ji, Difei Tang 외

One challenge for evaluating current sequence- or dialogue-level chatbots, such as Empathetic Open-domain Conversation Models, is to determine whether the chatbot performs in an emotionally consistent way. The most recen…

ChatbotDiversity