CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
Large Language Models (LLMs) have made significant progress in code generation, offering developers groundbreaking automated programming support. However, LLMs often generate code that is syntactically correct and even semantically plausible, but may not execute as expected or fulfill specified requirements. This phenomenon of hallucinations in the code domain has not been systematically explored. To advance the community's understanding and research on this issue, we introduce the concept of code hallucinations and propose a classification method for code hallucination based on execution verification. We categorize code hallucinations into four main types: mapping, naming, resource, and logic hallucinations, with each category further divided into different subcategories to understand and address the unique challenges faced by LLMs in code generation with finer granularity. Additionally, we present a dynamic detection algorithm called CodeHalu designed to detect and quantify code hallucinations. We also introduce the CodeHaluEval benchmark, which includes 8,883 samples from 699 tasks, to systematically and quantitatively evaluate code hallucinations. By evaluating 17 popular LLMs using this benchmark, we reveal significant differences in their accuracy and reliability in code generation, offering detailed insights for further improving the code generation capabilities of LLMs. The CodeHalu benchmark and code are publicly available at https://github.com/yuchen814/CodeHalu.
Code (1)
Tasks
Code GenerationHallucinationSimilar Papers 제목 키워드 기반
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
Despite their success, large language models (LLMs) face the critical challenge of hallucinations, generating plausible but incorrect content. While much research has focused on hallucinations in multiple modalities incl…
Code GenerationHallucinationLanguage ModelingLanguage Modelling+1Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
Recent technical breakthroughs in large language models (LLMs) have enabled them to fluently generate source code. Software developers often leverage both general-purpose and code-specialized LLMs to revise existing code…
Code GenerationHallucinationSurveyFrom Single to Multi: How LLMs Hallucinate in Multi-Document Summarization
Although many studies have investigated and reduced hallucinations in large language models (LLMs) for single-document tasks, research on hallucination in multi-document summarization (MDS) tasks remains largely unexplor…
Document SummarizationHallucinationMulti-Document SummarizationInvestigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
Large Language Models (LLMs) have achieved remarkable performance across a wide variety of natural language tasks. However, they have been shown to suffer from a critical limitation pertinent to 'hallucination' in their …
Abstractive Text SummarizationDialogue GenerationHallucinationLogical Reasoning+3CRABS: A syntactic-semantic pincer strategy for bounding LLM interpretation of Python notebooks
Recognizing the information flows and operations comprising data science and machine learning Python notebooks is critical for evaluating, reusing, and adapting notebooks for new tasks. Investigating a notebook via re-ex…
Zero-Shot Learning