paper-with-me

Papers

ChatEval: A Tool for Chatbot Evaluation

2019-06-01 · NAACL 2019 6 · Jo{\~a}o Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani, Lyle Ungar, Chris Callison-Burch

Open-domain dialog systems (i.e. chatbots) are difficult to evaluate. The current best practice for analyzing and comparing these dialog systems is the use of human judgments. However, the lack of standardization in evaluation procedures, and the fact that model parameters and code are rarely published hinder systematic human evaluation experiments. We introduce a unified framework for human evaluation of chatbots that augments existing tools and provides a web-based hub for researchers to share and compare their dialog systems. Researchers can submit their trained models to the ChatEval web interface and obtain comparisons with baselines and prior work. The evaluation code is open-source to ensure standardization and transparency. In addition, we introduce open-source baseline models and evaluation datasets. ChatEval can be found at https://chateval.org.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ChatbotOpen-Domain Dialog

Similar Papers 제목 키워드 기반

ChatEval: A Tool for the Systematic Evaluation of Chatbots

2018-11-01 · WS 2018 11 · Jo{\~a}o Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani 외
ChatbotText Generation

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

2023-08-14 · Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu 외

Text evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have explored LLMs' potential as alternative…

Text Generation

Three Ways of Using Large Language Models to Evaluate Chat

2023-08-12 · Ondřej Plátek, Vojtěch Hudeček, Patricia Schmidtová, Mateusz Lango 외

This paper describes the systems submitted by team6 for ChatEval, the DSTC 11 Track 4 competition. We present three different approaches to predicting turn-level qualities of chatbot responses based on large language mod…

Chatbot

Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop

2024-07-08 · Anum Afzal, Alexander Kowsik, Rajna Fani, Florian Matthes

Large Language Models have found application in various mundane and repetitive tasks including Human Resource (HR) support. We worked with the domain experts of SAP SE to develop an HR support chatbot as an efficient and…

ChatbotRetrieval

Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools

2024-08-03 · Jung In Park, Mahyar Abbasian, Iman Azimi, Dawn T. Bounds 외

Objective: This study aims to develop and validate an evaluation framework to ensure the safety and reliability of mental health chatbots, which are increasingly popular due to their accessibility, human-like interaction…

ChatbotLarge Language Model