paper-with-me

Papers

Benchmarking LLM powered Chatbots: Methods and Metrics

2023-08-08 · Debarag Banerjee, Pooja Singh, Arjun Avadhanam, Saksham Srivastava

Autonomous conversational agents, i.e. chatbots, are becoming an increasingly common mechanism for enterprises to provide support to customers and partners. In order to rate chatbots, especially ones powered by Generative AI tools like Large Language Models (LLMs) we need to be able to accurately assess their performance. This is where chatbot benchmarking becomes important. In this paper, we propose the use of a novel benchmark that we call the E2E (End to End) benchmark, and show how the E2E benchmark can be used to evaluate accuracy and usefulness of the answers provided by chatbots, especially ones powered by LLMs. We evaluate an example chatbot at different levels of sophistication based on both our E2E benchmark, as well as other available metrics commonly used in the state of art, and observe that the proposed benchmark show better results compared to others. In addition, while some metrics proved to be unpredictable, the metric associated with the E2E benchmark, which uses cosine similarity performed well in evaluating chatbots. The performance of our best models shows that there are several benefits of using the cosine similarity score as a metric in the E2E benchmark.

📄 PDF Abstract BibTeX arXiv:2308.04624

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingChatbot

Similar Papers 제목 키워드 기반

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI

2024-12-17 · Deep Bhatt, Surya Ayyagari, Anuruddh Mishra

Diagnostic errors in healthcare persist as a critical challenge, with increasing numbers of patients turning to online resources for health information. While AI-powered healthcare chatbots show promise, there exists no …

BenchmarkingChatbotDiagnostic

LLM-empowered Chatbots for Psychiatrist and Patient Simulation: Application and Evaluation

2023-05-23 · Siyuan Chen, Mengyue Wu, Kenny Q. Zhu, Kunyao Lan 외

Empowering chatbots in the field of mental health is receiving increasing amount of attention, while there still lacks exploration in developing and evaluating chatbots in psychiatric outpatient scenarios. In this work, …

ChatbotDiagnostic

Islamic Chatbots in the Age of Large Language Models

2025-12-31 · Muhammad Aurangzeb Ahmad arxiv

Large Language Models (LLMs) are rapidly transforming how communities access, interpret, and circulate knowledge, and religious communities are no exception. Chatbots powered by LLMs are beginning to reshape authority, p…

Assessing Empathy in Large Language Models with Real-World Physician-Patient Interactions

2024-05-26 · Man Luo, Christopher J. Warren, Lu Cheng, Haidar M. Abdul-Muhsin 외

The integration of Large Language Models (LLMs) into the healthcare domain has the potential to significantly enhance patient care and support through the development of empathetic, patient-facing chatbots. This study in…

A Machine Translation-Powered Chatbot for Public Administration

2022-06-01 · EAMT 2022 6 · Dimitra Anastasiou, Anders Ruge, Radu Ion, Svetlana Segărceanu 외

This paper is about a multilingual chatbot developed for public administration within the CEF funded project ENRICH4ALL. We argue for multi-lingual chatbots empowered through MT and discuss the integration of the CEF eTr…

ChatbotMachine TranslationTranslation