ChatEval: A Tool for the Systematic Evaluation of Chatbots
Code (0)
등록된 구현이 없습니다.
Tasks
ChatbotText GenerationSimilar Papers 제목 키워드 기반
ChatEval: A Tool for Chatbot Evaluation
Open-domain dialog systems (i.e. chatbots) are difficult to evaluate. The current best practice for analyzing and comparing these dialog systems is the use of human judgments. However, the lack of standardization in eval…
ChatbotOpen-Domain DialogChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Text evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have explored LLMs' potential as alternative…
Text GenerationBelief in Authority: Impact of Authority in Multi-Agent Evaluation Framework
Multi-agent systems utilizing large language models often assign authoritative roles to improve performance, yet the impact of authority bias on agent interactions remains underexplored. We present the first systematic a…
Conversational Process Modeling: Can Generative AI Empower Domain Experts in Creating and Redesigning Process Models?
AI-driven chatbots such as ChatGPT have caused a tremendous hype lately. For BPM applications, several applications for AI-driven chatbots have been identified to be promising to generate business value, including explan…
Systematic Literature ReviewOnline vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots
This paper explores the efficacy of online versus offline evaluation methods in assessing conversational chatbots, specifically comparing first-party direct interactions with third-party observational assessments. By ext…
BenchmarkingChatbot