paper-with-me

홈 › Papers

Data Contamination Through the Lens of Time

2023-10-16 · Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, Samuel Dooley

Recent claims about the impressive abilities of large language models (LLMs) are often supported by evaluating publicly available benchmarks. Since LLMs train on wide swaths of the internet, this practice raises concerns of data contamination, i.e., evaluating on examples that are explicitly or implicitly included in the training data. Data contamination remains notoriously challenging to measure and mitigate, even with partial attempts like controlled experimentation of training data, canary strings, or embedding similarities. In this work, we conduct the first thorough longitudinal analysis of data contamination in LLMs by using the natural experiment of training cutoffs in GPT models to look at benchmarks released over time. Specifically, we consider two code/mathematical problem-solving datasets, Codeforces and Project Euler, and find statistically significant trends among LLM pass rate vs. GitHub popularity and release date that provide strong evidence of contamination. By open-sourcing our dataset, raw results, and evaluation framework, our work paves the way for rigorous analyses of data contamination in modern models. We conclude with a discussion of best practices and future steps for publicly releasing benchmarks in the age of LLMs that train on webscale data.

📄 PDF Abstract BibTeX arXiv:2310.10628

Code (1)

abacusai/to-the-cutoff 공식 구현

Tasks

Mathematical Problem-Solving

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Residual Connection 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…

Similar Papers 제목 키워드 기반

A machine learning based approach to gravitational lens identification with the International LOFAR Telescope

2022-07-21 · S. Rezaei, J. P. McKean, M. Biehl, W. de Roo1 외

We present a novel machine learning based approach for detecting galaxy-scale gravitational lenses from interferometric data, specifically those taken with the International LOFAR Telescope (ILT), which is observing the …

ROBUST ESTIMATION VIA GENERATIVE ADVERSARIAL NETWORKS

2019-05-01 · ICLR 2019 5 · Chao GAO, jiyi LIU, Yuan YAO, Weizhi Zhu

Robust estimation under Huber's $\epsilon$-contamination model has become an important topic in statistics and theoretical computer science. Rate-optimal procedures such as Tukey's median and other estimators based on st…

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

2025-09-25 · Zirui He, Haiyan Zhao, Yingcong Li, Ali Payani 외 arxiv

Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memorized solutions appear as genuine reasoning. Existing detection methods l…

Label-free detection of Giardia lamblia cysts using a deep learning-enabled portable imaging flow cytometer

2020-07-12 · Zoltan Gorocs, David Baum, Fang Song, Kevin DeHaan 외

We report a field-portable and cost-effective imaging flow cytometer that uses deep learning to accurately detect Giardia lamblia cysts in water samples at a volumetric throughput of 100 mL/h. This flow cytometer uses le…

Robustness of Anomaly Detection Models for Industrial Control Systems under Training-Time Data Contamination

2026-08-24 · Mustafa Umut Ozbek, Taiwo Ojo, Pooria Madani, Khalil El-Khatib 외 arxiv

Machine-learning-based anomaly detection is increasingly used in industrial control systems (ICS), yet most studies assume that detector training data is trustworthy. In practice, training data may be corrupted through c…

Anomaly Detection