paper-with-me

Papers

Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests

2026-06-05 · Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori, Nontawat Charoenphakdee, Masashi Sugiyama, Takashi Ishida arxiv

A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing deceptive performance. This makes evaluation scores unreliable as measures of true task-solving ability. We propose CapCode, a framework for constructing coding datasets with randomized tests whose best achievable non-cheating performance is deliberately capped below one. This capped-performance design gives evaluation scores a clearer interpretation: scores substantially above the cap are implausible and therefore provide evidence of cheating. To prevent cheating, we propose CapReward, a reward design based on the CapCode principle to discourage optimization beyond the cap. Experiments across multiple datasets show that CapCode detects cheating while preserving performance ranking of models, and CapReward reduces cheating behavior, yielding models that better follow the intended task specification.

📄 PDF Abstract BibTeX arXiv:2606.07379

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Examining Monitoring System: Detecting Abnormal Behavior In Online Examinations

2024-02-19 · Dinh An Ngo, Thanh Dat Nguyen, Thi Le Chi Dang, Huy Hoan Le 외

Cheating in online exams has become a prevalent issue over the past decade, especially during the COVID-19 pandemic. To address this issue of academic dishonesty, our "Exam Monitoring System: Detecting Abnormal Behavior …

Decision Making

A Video-based Detector for Suspicious Activity in Examination with OpenPose

2023-07-21 · Reuben Moyo, Stanley Ndebvu, Michael Zimba, Jimmy Mbelwa

Examinations are a crucial part of the learning process, and academic institutions invest significant resources into maintaining their integrity by preventing cheating from students or facilitators. However, cheating has…

Fairness

GAN-Aimbots: Using Machine Learning for Cheating in First Person Shooters

2022-05-14 · Anssi Kanervisto, Tomi Kinnunen, Ville Hautamäki

Playing games with cheaters is not fun, and in a multi-billion-dollar video game industry with hundreds of millions of players, game developers aim to improve the security and, consequently, the user experience of their …

BIG-bench Machine Learning

Human-in-the-Loop AI for Cheating Ring Detection

2024-03-18 · Yong-Siang Shih, Manqian Liao, Ruidong Liu, Mirza Basim Baig

Online exams have become popular in recent years due to their accessibility. However, some concerns have been raised about the security of the online exams, particularly in the context of professional cheating services a…

Fairness

Applying IRT to Distinguish Between Human and Generative AI Responses to Multiple-Choice Assessments

2024-11-28 · Alona Strugatski, Giora Alexandron

Generative AI is transforming the educational landscape, raising significant concerns about cheating. Despite the widespread use of multiple-choice questions in assessments, the detection of AI cheating in MCQ-based test…

Multiple-choice