paper-with-me

홈 › Papers

Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories

2026-06-23 · Arsham Khosravani, Audris Mockus arxiv

Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence poorly understood. We introduce a multi-layered detection framework that integrates configuration-file scanning, commit-message analysis, author-identity matching, and bot-signature lookup across World of Code (180M+ Git repositories), classifying agent traces into four behavioral types. No single method captures more than a fraction of activity: multi-method detection identifies 850,157 Claude Code commits in one snapshot, of which bot-account lookup_the signal most adoption studies rely on_recovers only 28,154 (3.3%), a 30x relative-recall gap, so single-signal prevalence estimates are biased low by at least this factor. Every detection pattern is hand-validated (495 labels) with per-cell precision and Wilson confidence intervals. Across snapshots from December 2024 to April 2026, commit-attributed agents generate over 320,000 commits per month; Claude Code leads (886,122 commits across 17,295 projects) and dominates silent, configuration-file-only adoption (21,078 projects). Compared against an independent pull-request census (AIDev), the two channels capture nearly disjoint agent populations_a PR census misses 79% of commit-detected Claude Code adopters and essentially all Codex adopters_and different kinds of work: PR-deployed cloud agents (Codex, Cursor) surface as feature work, while commit-deployed in-editor agents (Claude Code, OpenHands, Aider) surface as maintenance. The observed work profile follows deployment and detection mode rather than the tool itself, so no single channel is representative.

📄 PDF Abstract BibTeX arXiv:2606.24429

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

2025-12-20 · Tue Le, Minh V. T. Thai, Dung Nguyen Manh, Huy Phan Nhat 외 arxiv

Existing benchmarks for AI coding agents focus on isolated, single-issue tasks such as fixing a bug or adding a small feature. However, real-world software engineering is a long-horizon endeavor: developers interpret hig…

MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research

2025-05-26 · Hui Chen, Miao Xiong, Yujie Lu, Wei Han 외

Recent advancements in AI agents have demonstrated their growing potential to drive and support scientific discovery. In this work, we introduce MLR-Bench, a comprehensive benchmark for evaluating AI agents on open-ended…

scientific discovery

SERA: Soft-Verified Efficient Repository Agents

2026-01-28 · Ethan Shen, Daniel Tormoen, Saurabh Shah, Ali Farhadi 외 arxiv

Open-weight coding agents should hold a fundamental advantage over closed-source systems because they can specialize to private codebases, encoding repository-specific information directly in their weights. Yet the cost …

Reinforcement Learning

EvilGenie: A Reward Hacking Benchmark

2025-11-26 · Jonathan Gabor, Jayson Lynch, Jonathan Rosenfeld arxiv

We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in which agents can easily reward hack, such as by hardcoding test cases or …

VisCoder2: Building Multi-Language Visualization Coding Agents

2025-10-24 · Yuansheng Ni, Songcheng Cai, Xiangchao Chen, Jiarong Liang 외 arxiv

Large language models (LLMs) have recently enabled coding agents capable of generating, executing, and revising visualization code. However, existing models often fail in practical workflows due to limited language cover…