paper-with-me

홈 › Papers

Decontamination of the scientific literature

2022-10-28 · Guillaume Cabanac

Research misconduct and frauds pollute the scientific literature. Honest errors and malevolent data fabrication, image manipulation, journal hijacking, and plagiarism passed peer review unnoticed. Problematic papers deceive readers, authors citing them, and AI-powered literature-based discovery. Flagship publishers accepted hundreds flawed papers despite claiming to enforce peer review. This application ambitions to decontaminate the scientific literature using curative and preventive actions.

📄 PDF Abstract BibTeX arXiv:2210.15912

Code (0)

등록된 구현이 없습니다.

Tasks

Image Manipulation

Similar Papers 제목 키워드 기반

Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

2023-11-08 · Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E. Gonzalez 외

Large language models are increasingly trained on all the data ever produced by humans. Many have raised concerns about the trustworthiness of public benchmarks due to potential contamination in pre-training or fine-tuni…

HumanEvalMMLU

Improving Human Judgments by Decontaminating Sequential Dependencies

2010-12-01 · NeurIPS 2010 12 · Michael C. Mozer, Harold Pashler, Matthew Wilder, Robert V. Lindsey 외

For over half a century, psychologists have been struck by how poor people are at expressing their internal sensations, impressions, and evaluations via rating scales. When individuals make judgments, they are incapable …

Analysis of bacterial population growth using extended logistic growth model with distributed delay

2018-07-21

In the present work, we develop a delayed Logistic growth model to study the effects of decontamination on the bacterial population in the ambient environment. Using the linear stability analysis, we study different case…

Uncertainty-based Debiasing and Unlearning for Decontamination

2026-06-22 · Guangzhi Sun, Xiao Zhan, Mark Gales arxiv

Benchmark-based evaluation is the dominant paradigm for assessing large language model (LLM) capabilities, yet data contamination inflates reported performance and undermines fair comparison. Existing decontamination met…

Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation

2024-06-20 · Qin Zhu, Qingyuan Cheng, Runyu Peng, Xiaonan Li 외

The training process of large language models (LLMs) often involves varying degrees of test data contamination. Although current LLMs are achieving increasingly better performance on various benchmarks, their performance…

GSM8KLanguage Model EvaluationLanguage ModelingLanguage Modelling+2