paper-with-me

홈 › Papers

Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models

2024-03-06 · Martin Riddell, Ansong Ni, Arman Cohan

While large language models have achieved remarkable performance on various code generation benchmarks, there have been growing concerns regarding potential contamination of these benchmarks as they may be leaked into pretraining and finetuning data. While recent work has investigated contamination in natural language generation and understanding tasks, there has been less extensive research into how data contamination impacts the evaluation of code generation, which is critical for understanding the robustness and reliability of LLMs in programming contexts. In this work, we perform a comprehensive study of data contamination of popular code generation benchmarks, and precisely quantify their overlap with pretraining corpus through both surface-level and semantic-level matching. In our experiments, we show that there are substantial overlap between popular code generation benchmarks and open training corpus, and models perform significantly better on the subset of the benchmarks where similar solutions are seen during training. We also conduct extensive analysis on the factors that affects model memorization and generalization, such as model size, problem difficulty, and question length. We release all resulting files from our matching pipeline for future research.

📄 PDF Abstract BibTeX arXiv:2403.04811

Code (1)

yale-nlp/code-llm-contamination 공식 구현

Tasks

Code GenerationMemorizationText Generation

Similar Papers 제목 키워드 기반

Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination

2025-03-06 · Simin Chen, Pranav Pusarla, Baishakhi Ray

The rapid evolution of code largelanguage models underscores the need for effective and transparent benchmarking of their reasoning capabilities. However, the current benchmarking approach heavily depends on publicly ava…

Benchmarking

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

2026-06-18 · Maria Ivanova, Pavel Zadorozhny, Rodion Levichev, Ivan Petrov 외 arxiv

LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on code-generation tasks. By curating competitive programming problems, constantly adding fresh problems to t…

Code Generation

CodeReviewQA: The Code Review Comprehension Assessment for Large Language Models

2025-03-20 · Hong Yi Lin, Chunhua Liu, Haoyu Gao, Patanamon Thongtanunam 외

State-of-the-art large language models (LLMs) have demonstrated impressive code generation capabilities but struggle with real-world software engineering tasks, such as revising source code to address code reviews, hinde…

Code GenerationMultiple-choiceText Matching

M2G-Eval: Enhancing and Evaluating Multi-granularity Multilingual Code Generation

2025-12-27 · Fanglin Xu, Wei Zhang, Jian Yang, Guo Chen 외 arxiv

The rapid advancement of code large language models (LLMs) has sparked significant research interest in systematically evaluating their code generation capabilities, yet existing benchmarks predominantly assess models at…

Code Generation

NotSoTiny: A Large, Living Benchmark for RTL Code Generation

2025-12-23 · Razine Moundir Ghorab, Emanuele Parisi, Cristian Gutierrez, Miquel Alberti-Binimelis 외 arxiv

LLMs have shown early promise in generating RTL code, yet evaluating their capabilities in realistic setups remains a challenge. So far, RTL benchmarks have been limited in scale, skewed toward trivial designs, offering …

Code Generation