paper-with-me

홈 › Papers

Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

2024-10-13 · Nan Jiang, Qi Li, Lin Tan, Tianyi Zhang

Despite their success, large language models (LLMs) face the critical challenge of hallucinations, generating plausible but incorrect content. While much research has focused on hallucinations in multiple modalities including images and natural language text, less attention has been given to hallucinations in source code, which leads to incorrect and vulnerable code that causes significant financial loss. To pave the way for research in LLMs' hallucinations in code, we introduce Collu-Bench, a benchmark for predicting code hallucinations of LLMs across code generation (CG) and automated program repair (APR) tasks. Collu-Bench includes 13,234 code hallucination instances collected from five datasets and 11 diverse LLMs, ranging from open-source models to commercial ones. To better understand and predict code hallucinations, Collu-Bench provides detailed features such as the per-step log probabilities of LLMs' output, token types, and the execution feedback of LLMs' generated code for in-depth analysis. In addition, we conduct experiments to predict hallucination on Collu-Bench, using both traditional machine learning techniques and neural networks, which achieves 22.03 -- 33.15% accuracy. Our experiments draw insightful findings of code hallucination patterns, reveal the challenge of accurately localizing LLMs' hallucinations, and highlight the need for more sophisticated techniques.

📄 PDF Abstract BibTeX arXiv:2410.09997

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationHallucinationLanguage ModelingLanguage ModellingProgram Repair

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Collusion Detection with Graph Neural Networks

2024-10-09 · Lucas Gomes, Jannis Kueck, Mara Mattes, Martin Spindler 외

Collusion is a complex phenomenon in which companies secretly collaborate to engage in fraudulent practices. This paper presents an innovative methodology for detecting and predicting collusion patterns in different nati…

FairnessTransfer LearningZero-Shot Learning

Audit the Whisper: Detecting Steganographic Collusion in Multi-Agent LLMs

2025-10-05 · Om Tailor arxiv

Multi-agent deployments of large language models (LLMs) are increasingly embedded in market, allocation, and governance workflows, yet covert coordination among agents can silently erode trust and social welfare. Existin…

Testing Models of Strategic Uncertainty: Equilibrium Selection in Repeated Games

2021-01-14 · Emanuel Vespa, Taylor Weidman, Alistair J. Wilson

In repeated-game applications where both the collusive and non-collusive outcomes can be supported as equilibria, researchers must resolve underlying selection questions if theory will be used to understand counterfactua…

counterfactual

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

2024-05-08 · CVPR 2024 1 · Prannay Kaul, Zhizhong Li, Hao Yang, Yonatan Dukler 외

Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term "Type I hallucinations". Instead…

AttributeData AugmentationFormHallucination+1

Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding

2024-10-17 · Kyungmin Min, Minbeom Kim, Kang-il Lee, Dongryeol Lee 외

Large Vision-Language Models (LVLMs) demonstrate impressive capabilities in generating detailed and coherent responses from visual inputs. However, they are prone to generate hallucinations due to an over-reliance on lan…

HallucinationObject HallucinationPOS