paper-with-me

Papers

Humans are Still Better than ChatGPT: Case of the IEEEXtreme Competition

2023-05-10 · Anis Koubaa, Basit Qureshi, Adel Ammar, Zahid Khan, Wadii Boulila, Lahouari Ghouti

Since the release of ChatGPT, numerous studies have highlighted the remarkable performance of ChatGPT, which often rivals or even surpasses human capabilities in various tasks and domains. However, this paper presents a contrasting perspective by demonstrating an instance where human performance excels in typical tasks suited for ChatGPT, specifically in the domain of computer programming. We utilize the IEEExtreme Challenge competition as a benchmark, a prestigious, annual international programming contest encompassing a wide range of problems with different complexities. To conduct a thorough evaluation, we selected and executed a diverse set of 102 challenges, drawn from five distinct IEEExtreme editions, using three major programming languages: Python, Java, and C++. Our empirical analysis provides evidence that contrary to popular belief, human programmers maintain a competitive edge over ChatGPT in certain aspects of problem-solving within the programming context. In fact, we found that the average score obtained by ChatGPT on the set of IEEExtreme programming problems is 3.9 to 5.8 times lower than the average human score, depending on the programming language. This paper elaborates on these findings, offering critical insights into the limitations and potential areas of improvement for AI-based language models like ChatGPT.

📄 PDF Abstract BibTeX arXiv:2305.06934

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring ChatGPT's Empathic Abilities

2023-08-07 · Kristina Schaaff, Caroline Reinig, Tim Schlippe

Empathy is often understood as the ability to share and understand another individual's state of mind or emotion. With the increasing use of chatbots in various domains, e.g., children seeking help with homework, individ…

Chatbot

Language in Vivo vs. in Silico: Size Matters but Larger Language Models Still Do Not Comprehend Language on a Par with Humans

2024-04-23 · Vittoria Dentella, Fritz Guenther, Evelina Leivada

Understanding the limits of language is a prerequisite for Large Language Models (LLMs) to act as theories of natural language. LLM performance in some language tasks presents both quantitative and qualitative difference…

Non-native speakers of English or ChatGPT: Who thinks better?

2024-11-30 · Mohammed Q. Shormani

This study sets out to answer one major question: Who thinks better, non-native speakers of English or ChatGPT?, providing evidence from processing and interpreting center-embedding English constructions that human brain…

Sentence

Current LLMs still cannot 'talk much' about grammar modules: Evidence from syntax

2026-03-20 · Mohammed Q. Shormani, Yehia A. AlSohbani arxiv

We aim to examine the extent to which Large Language Models (LLMs) can 'talk much' about grammar modules, providing evidence from syntax core properties translated by ChatGPT into Arabic. We collected 44 terms from gener…

Benchmarking ChatGPT-4 on ACR Radiation Oncology In-Training (TXIT) Exam and Red Journal Gray Zone Cases: Potentials and Challenges for AI-Assisted Medical Education and Decision Making in Radiation Oncology

2023-04-24 · Yixing Huang, Ahmed Gomaa, Sabine Semrau, Marlen Haderlein 외

The potential of large language models in medicine for education and decision making purposes has been demonstrated as they achieve decent scores on medical exams such as the United States Medical Licensing Exam (USMLE) …

BenchmarkingDecision MakingHallucinationMedQA+1