paper-with-me

홈 › Papers

MultiAIGCD: A Comprehensive dataset for AI Generated Code Detection Covering Multiple Languages, Models,Prompts, and Scenarios

2025-07-29 · Basak Demirok, Mucahid Kutlu, Selin Mergen arxiv

As large language models (LLMs) rapidly advance, their role in code generation has expanded significantly. While this offers streamlined development, it also creates concerns in areas like education and job interviews. Consequently, developing robust systems to detect AI-generated code is imperative to maintain academic integrity and ensure fairness in hiring processes. In this study, we introduce MultiAIGCD, a dataset for AI-generated code detection for Python, Java, and Go. From the CodeNet dataset's problem definitions and human-authored codes, we generate several code samples in Java, Python, and Go with six different LLMs and three different prompts. This generation process covered three key usage scenarios: (i) generating code from problem descriptions, (ii) fixing runtime errors in human-written code, and (iii) correcting incorrect outputs. Overall, MultiAIGCD consists of 121,271 AI-generated and 32,148 human-written code snippets. We also benchmark three state-of-the-art AI-generated code detection models and assess their performance in various test scenarios such as cross-model and cross-language. We share our dataset and codes to support research in this field.

📄 PDF Abstract BibTeX arXiv:2507.21693

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT

2024-06-13 · Zhen Tao, Yanfang Chen, Dinghao Xi, Zhiyu Li 외

The increasing prevalence of large language models (LLMs) has significantly advanced text generation, but the human-like quality of LLM outputs presents major challenges in reliably distinguishing between human-authored …

BenchmarkingLLM-generated Text DetectionQuestion AnsweringText Detection+1

Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition

2025-01-11 · Yuankun Xie, Xiaopeng Wang, Zhiyong Wang, Ruibo Fu 외

Current research in audio deepfake detection is gradually transitioning from binary classification to multi-class tasks, referred as audio deepfake source tracing task. However, existing studies on source tracing conside…

Audio Deepfake DetectionBinary ClassificationDeepFake DetectionFace Swapping

AICD Bench: A Challenging Benchmark for AI-Generated Code Detection

2026-02-02 · Daniil Orel, Dilshod Azizov, Indraneil Paul, Yuxia Wang 외 arxiv

Large language models (LLMs) are increasingly capable of generating functional source code, raising concerns about authorship, accountability, and security. While detecting AI-generated code is critical, existing dataset…

Binary Classification

On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection

2024-10-31 · Xiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li 외

Large numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection al…

Video Forensics

Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities

2024-07-29 · Federico Cocchi, Marcella Cornia, Lorenzo Baraldi, Alessandro Nicolosi 외

Discerning between authentic content and that generated by advanced AI methods has become increasingly challenging. While previous research primarily addresses the detection of fake faces, the identification of generated…

Contrastive LearningDeepFake DetectionFace SwappingImage to text