paper-with-me

홈 › Papers

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems

2026-05-11 · Zonglin Yang, Xingtong Liu, Xinyan Xu arxiv

AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated. We introduce SCIINTEGRITY-BENCH, the first benchmark designed around a dilemmatic evaluation paradigm: each of its 33 scenarios across 11 trap categories is constructed so that honest acknowledgment of failure is the only correct response, while task completion requires misconduct. Across 231 evaluation runs spanning 7 state-of-the-art LLMs, the overall integrity problem rate reaches 34.2%, and no model achieves zero failures. Most strikingly, across missing-data scenarios, all seven models generate synthetic data rather than acknowledging infeasibility, differing only in whether they disclose the substitution. A further prompt ablation study separates two drivers: removing explicit completion pressure sharply reduces undisclosed fabrication from 20.6% to 3.2%, while the underlying synthesis rate remains unchanged, revealing an intrinsic completion bias that persists independent of prompt-level instructions. These findings point to the absence of honest refusal as a trained disposition as the primary driver of observed failures. We release SCIINTEGRITY-BENCH at https://github.com/liuxingtong/Sci-Integrity-Bench.

📄 PDF Abstract BibTeX arXiv:2605.10246

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DoPE: Decoy Oriented Perturbation Encapsulation Human-Readable, AI-Hostile Documents for Academic Integrity

2026-01-18 · Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi, Yash Shah 외 arxiv

Multimodal Large Language Models (MLLMs) can directly consume exam documents, threatening conventional assessments and academic integrity. We present DoPE (Decoy-Oriented Perturbation Encapsulation), a document-layer def…

Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges

2026-04-13 · Hatem M. El-boghdadi, Toqeer Ali Syed, Ali Akarma, Qamar Wali arxiv

The rapid adoption of AI tools such as ChatGPT has significantly transformed academic practices, offering considerable benefits for both students and faculty in computing disciplines. These tools have been shown to enhan…

DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys

2026-01-13 · Guo-Biao Zhang, Ding-Yuan Liu, Da-Yi Wu, Tian Lan 외 arxiv

The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys. Most existing benchmarks first const…

IntegrityAI at GenAI Detection Task 2: Detecting Machine-Generated Academic Essays in English and Arabic Using ELECTRA and Stylometry

2025-01-07 · Mohammad AL-Smadi

Recent research has investigated the problem of detecting machine-generated essays for academic purposes. To address this challenge, this research utilizes pre-trained, transformer-based models fine-tuned on Arabic and E…

Task 2

AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images

2026-04-30 · Bo Zhang, Tzu-Yen Ma, Zichen Tang, Junpeng Ding 외 arxiv

We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) Domain-Specific Complexity: covering seve…