paper-with-me

홈 › Papers

ChatEval: A Tool for the Systematic Evaluation of Chatbots

2018-11-01 · WS 2018 11 · Jo{\~a}o Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani, Lyle Ungar, Chris Callison-Burch
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ChatbotText Generation

Similar Papers 제목 키워드 기반

ChatEval: A Tool for Chatbot Evaluation

2019-06-01 · NAACL 2019 6 · Jo{\~a}o Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani 외

Open-domain dialog systems (i.e. chatbots) are difficult to evaluate. The current best practice for analyzing and comparing these dialog systems is the use of human judgments. However, the lack of standardization in eval…

ChatbotOpen-Domain Dialog

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

2023-08-14 · Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu 외

Text evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have explored LLMs' potential as alternative…

Text Generation

Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework

2026-01-08 · Junhyuk Choi, Jeongyoun Kwon, Heeju Kim, Haeun Cho 외 arxiv

Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexplored. We present the first systematic a…

Conversational Process Modeling: Can Generative AI Empower Domain Experts in Creating and Redesigning Process Models?

2023-04-19 · Nataliia Klievtsova, Janik-Vasily Benzin, Timotheus Kampik, Juergen Mangler 외

AI-driven chatbots such as ChatGPT have caused a tremendous hype lately. For BPM applications, several applications for AI-driven chatbots have been identified to be promising to generate business value, including explan…

Systematic Literature Review

Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots

2024-09-12 · Ekaterina Svikhnushina, Pearl Pu

This paper explores the efficacy of online versus offline evaluation methods in assessing conversational chatbots, specifically comparing first-party direct interactions with third-party observational assessments. By ext…

BenchmarkingChatbot