paper-with-me

홈 › Papers

An Empirical Study of Security Calibration in Large Language Models for Code

2026-06-30 · Mohammed Latif Siddiq, Md. Nafiu Rahman, Joanna C. S. Santos arxiv

Large Language Models (LLMs) are rapidly transforming software development, yet their use in security-critical contexts raises a key question: do models know when their generated code is insecure? This property, known as calibration, measures whether a model's confidence aligns with the true correctness of its outputs. We present the first large-scale empirical study of security calibration in LLM-generated code. We evaluate GPT-4o-mini, Gemini-2.0-Flash, and Qwen3-Coder-Next across multiple temperature settings on two complementary benchmarks: self-contained security tasks and multi-language repository-level contexts. Our results suggest that overconfidence is prevalent across the evaluated LLMs. Functional calibration is consistently worse than security calibration, suggesting that models estimate security outcomes more reliably than functional correctness, potentially because functional correctness depends on complex execution behavior. We also examine whether calibration-guided automated repair can help remediate vulnerabilities in LLM-generated code, finding only limited improvements while frequently introducing functional regressions. Moreover, we study different mitigation strategies for reducing False Trust, where models assign high confidence to vulnerable code. The results show that although architectural gating improves calibration on controlled benchmarks, calibration deteriorates in realistic repository-level settings, increasing the risk of high-confidence vulnerable outputs.

📄 PDF Abstract BibTeX arXiv:2606.31159

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts

2026-01-30 · Mona Rajhans arxiv

Artificial intelligence (AI) copilots are increasingly integrated into enterprise cybersecurity platforms to assist analysts in threat detection, triage, and remediation. However, the effectiveness of these systems depen…

The Dunning-Kruger Effect in Large Language Models: An Empirical Study of Confidence Calibration

2026-02-12 · Sudipta Ghosh, Mrityunjoy Panday arxiv

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet their ability to accurately assess their own confidence remains poorly understood. We present an empirical study investigat…

Confidence Calibration in Large Language Model-Based Entity Matching

2025-09-23 · Iris Kamsteeg, Juan Cardenas-Cartagena, Floris van Beers, Gineke ten Holt 외 arxiv

This research aims to explore the intersection of Large Language Models and confidence calibration in Entity Matching. To this end, we perform an empirical study to compare baseline RoBERTa confidences for an Entity Matc…

On Calibration of Large Language Models: From Response To Capability

2026-02-14 · Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin, Yun-Nung Chen 외 arxiv

Large language models (LLMs) are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use. Prior work on LLM calibration largely focuses on response-level confid…

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

2026-06-29 · Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li 외 arxiv

Calibration evaluates whether a model confidence aligns with its empirical accuracy. Existing studies often compare the calibration of different large language models using global calibration metrics such as Expected Cal…