paper-with-me

홈 › Papers

The Dunning-Kruger Effect in Large Language Models: An Empirical Study of Confidence Calibration

2026-02-12 · Sudipta Ghosh, Mrityunjoy Panday arxiv

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet their ability to accurately assess their own confidence remains poorly understood. We present an empirical study investigating whether LLMs exhibit patterns reminiscent of the Dunning-Kruger effect -- a cognitive bias where individuals with limited competence tend to overestimate their abilities. We evaluate four state-of-the-art models (Claude Haiku 4.5, Gemini 2.5 Pro, Gemini 2.5 Flash, and Kimi K2) across four benchmark datasets totaling 24,000 experimental trials. Our results reveal striking calibration differences: Kimi K2 exhibits severe overconfidence with an Expected Calibration Error (ECE) of 0.726 despite only 23.3% accuracy, while Claude Haiku 4.5 achieves the best calibration (ECE = 0.122) with 75.4% accuracy. These findings demonstrate that poorly performing models display markedly higher overconfidence -- a pattern analogous to the Dunning-Kruger effect in human cognition. We discuss implications for safe deployment of LLMs in high-stakes applications.

📄 PDF Abstract BibTeX arXiv:2603.09985

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond the Steeper Curve: AI-Mediated Metacognitive Decoupling and the Limits of the Dunning-Kruger Metaphor

2026-03-31 · Christopher Koch arxiv

The common claim that generative AI simply amplifies the Dunning-Kruger effect is too coarse to capture the available evidence. The clearest findings instead suggest that large language model (LLM) use can improve observ…

Do Code Models Suffer from the Dunning-Kruger Effect?

2025-10-06 · Mukul Singh, Somya Chatterjee, Arjun Radhakrishna, Sumit Gulwani arxiv

As artificial intelligence systems increasingly collaborate with humans in creative and technical domains, questions arise about the cognitive boundaries and biases that shape our shared agency. This paper investigates t…

The Confidence-Competence Gap in Large Language Models: A Cognitive Study

2023-09-28 · Aniket Kumar Singh, Suman Devkota, Bishal Lamichhane, Uttam Dhakal 외

Large Language Models (LLMs) have acquired ubiquitous attention for their performances across diverse domains. Our study here searches through LLMs' cognitive abilities and confidence dynamics. We dive deep into understa…

Scaling Truth: The Confidence Paradox in AI Fact-Checking

2025-09-10 · Ihsan A. Qazi, Zohaib Khan, Abdullah Ghani, Agha A. Raza 외 arxiv

The rise of misinformation underscores the need for scalable and reliable fact-checking solutions. Large language models (LLMs) hold promise in automating fact verification, yet their effectiveness across global contexts…

Fact Verification

ChainBuddy: An AI Agent System for Generating LLM Pipelines

2024-09-20 · Jingyue Zhang, Ian Arawjo

As large language models (LLMs) advance, their potential applications have grown significantly. However, it remains difficult to evaluate LLM behavior on user-defined tasks and craft effective pipelines to do so. Many us…

AI Agent