paper-with-me

Papers

The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models

2025-01-16 · Jonathan Katzy, Razvan Mihai Popescu, Arie van Deursen, Maliheh Izadi

The recent rise in the popularity of large language models has spurred the development of extensive code datasets needed to train them. This has left limited code available for collection and use in the downstream investigation of specific behaviors, or evaluation of large language models without suffering from data contamination. To address this problem, we release The Heap, a large multilingual dataset covering 57 programming languages that has been deduplicated with respect to other open datasets of code, enabling researchers to conduct fair evaluations of large language models without significant data cleaning overhead.

📄 PDF Abstract BibTeX arXiv:2501.09653

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Contamination Report for Multilingual Benchmarks

2024-10-21 · Sanchit Ahuja, Varun Gumma, Sunayana Sitaram

Benchmark contamination refers to the presence of test datasets in Large Language Model (LLM) pre-training or post-training data. Contamination can lead to inflated scores on benchmarks, compromising evaluation results a…

Language ModelingLanguage ModellingLarge Language Model

Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora

2026-01-21 · Chaymaa Abbas, Nour Shamaa, Mariette Awad arxiv

Data contamination undermines the validity of Large Language Model evaluation by enabling models to rely on memorized benchmark content rather than true generalization. While prior work has proposed contamination detecti…

A Novel Heap-based Pilot Assignment for Full Duplex Cell-Free Massive MIMO with Zero-Forcing

2020-07-08 · Hieu V. Nguyen, Van-Dinh Nguyen, Octavia A. Dobre, Shree Krishna Sharma 외

This paper investigates the combined benefits of full-duplex (FD) and cell-free massive multiple-input multipleoutput (CF-mMIMO), where a large number of distributed access points (APs) having FD capability simultaneousl…

Robust Design

Data Contamination Can Cross Language Barriers

2024-06-19 · Feng Yao, Yufan Zhuang, Zihao Sun, Sunan Xu 외

The opacity in developing large language models (LLMs) is raising growing concerns about the potential contamination of public benchmarks in the pre-training data. Existing contamination detection methods are typically b…

Memorization

LNE-Blocking: An Efficient Framework for Contamination Mitigation Evaluation on Large Language Models

2025-09-18 · Ruijie Hou, Yueyang Jiao, Hanxu Hu, Yingming Li 외 arxiv

The problem of data contamination is now almost inevitable during the development of large language models (LLMs), with the training data commonly integrating those evaluation benchmarks even unintentionally. This proble…