ChatMatch: Evaluating Chatbots by Autonomous Chat Tournaments
Existing automatic evaluation systems of chatbots mostly rely on static chat scripts as ground truth, which is hard to obtain, and requires access to the models of the bots as a form of “white-box testing”. Interactive evaluation mitigates this problem but requires human involvement. In our work, we propose an interactive chatbot evaluation framework in which chatbots compete with each other like in a sports tournament, using flexible scoring metrics. This framework can efficiently rank chatbots independently from their model architectures and the domains for which they are trained.
Code (1)
Tasks
ChatbotSimilar Papers 제목 키워드 기반
Benchmarking LLM powered Chatbots: Methods and Metrics
Autonomous conversational agents, i.e. chatbots, are becoming an increasingly common mechanism for enterprises to provide support to customers and partners. In order to rate chatbots, especially ones powered by Generativ…
BenchmarkingChatbotLLM-empowered Chatbots for Psychiatrist and Patient Simulation: Application and Evaluation
Empowering chatbots in the field of mental health is receiving increasing amount of attention, while there still lacks exploration in developing and evaluating chatbots in psychiatric outpatient scenarios. In this work, …
ChatbotDiagnosticIf I Hear You Correctly: Building and Evaluating Interview Chatbots with Active Listening Skills
Interview chatbots engage users in a text-based conversation to draw out their views and opinions. It is, however, challenging to build effective interview chatbots that can handle user free-text responses to open-ended …
ChatbotEvaluating Chatbots to Promote Users' Trust -- Practices and Open Problems
Chatbots, the common moniker for collaborative assistants, are Artificial Intelligence (AI) software that enables people to naturally interact with them to get tasks done. Although chatbots have been studied since the da…
ChatbotLanguage ModelingLanguage ModellingLarge Language ModelEvaluator for Emotionally Consistent Chatbots
One challenge for evaluating current sequence- or dialogue-level chatbots, such as Empathetic Open-domain Conversation Models, is to determine whether the chatbot performs in an emotionally consistent way. The most recen…
ChatbotDiversity