paper-with-me

홈 › Papers

Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard

2023-05-30 · Vagelis Plevris, George Papazafeiropoulos, Alejandro Jiménez Rios

A comparison between three chatbots which are based on large language models, namely ChatGPT-3.5, ChatGPT-4 and Google Bard is presented, focusing on their ability to give correct answers to mathematics and logic problems. In particular, we check their ability to Understand the problem at hand; Apply appropriate algorithms or methods for its solution; and Generate a coherent response and a correct answer. We use 30 questions that are clear, without any ambiguities, fully described with plain text only, and have a unique, well defined correct answer. The questions are divided into two sets of 15 each. The questions of Set A are 15 "Original" problems that cannot be found online, while Set B contains 15 "Published" problems that one can find online, usually with their solution. Each question is posed three times to each chatbot. The answers are recorded and discussed, highlighting their strengths and weaknesses. It has been found that for straightforward arithmetic, algebraic expressions, or basic logic puzzles, chatbots may provide accurate solutions, although not in every attempt. However, for more complex mathematical problems or advanced logic tasks, their answers, although written in a usually "convincing" way, may not be reliable. Consistency is also an issue, as many times a chatbot will provide conflicting answers when given the same question more than once. A comparative quantitative evaluation of the three chatbots is made through scoring their final answers based on correctness. It was found that ChatGPT-4 outperforms ChatGPT-3.5 in both sets of questions. Bard comes third in the original questions of Set A, behind the other two chatbots, while it has the best performance (first place) in the published questions of Set B. This is probably because Bard has direct access to the internet, in contrast to ChatGPT chatbots which do not have any communication with the outside world.

📄 PDF Abstract BibTeX arXiv:2305.18618

Code (0)

등록된 구현이 없습니다.

Tasks

ChatbotMath

Similar Papers 제목 키워드 기반

Developing Effective Educational Chatbots with ChatGPT prompts: Insights from Preliminary Tests in a Case Study on Social Media Literacy (with appendix)

2023-06-18 · Cansu Koyuturk, Mona Yavari, Emily Theophilou, Sathya Bursic 외

Educational chatbots come with a promise of interactive and personalized learning experiences, yet their development has been limited by the restricted free interaction capabilities of available platforms and the difficu…

ChatbotZero-Shot Learning

Analyzing Large language models chatbots: An experimental approach using a probability test

2024-07-10 · Melise Peruchini, Julio Monteiro Teixeira

This study consists of qualitative empirical research, conducted through exploratory tests with two different Large Language Models (LLMs) chatbots: ChatGPT and Gemini. The methodological procedure involved exploratory t…

ChatbotLogical Reasoning

Evaluating Chatbots to Promote Users' Trust -- Practices and Open Problems

2023-09-09 · Biplav Srivastava, Kausik Lakkaraju, Tarmo Koppel, Vignesh Narayanan 외

Chatbots, the common moniker for collaborative assistants, are Artificial Intelligence (AI) software that enables people to naturally interact with them to get tasks done. Although chatbots have been studied since the da…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model

Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task

2024-06-28 · Sion Yoon, Tae Eun Kim, Yoo Jung Oh

The dynamics of human-AI communication have been reshaped by language models such as ChatGPT. However, extant research has primarily focused on dyadic communication, leaving much to be explored regarding the dynamics of …

ChatbotLanguage ModelingLanguage Modelling

Mental Health Assessment for the Chatbots

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Previous researches on dialogue system assessment usually focus on the quality evaluation (e.g. fluency, relevance, etc) of responses generated by the chatbots, which are local and technical metrics. For a chatbot which …

Chatbot