paper-with-me

Papers

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

2026-08-24 · Deyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na arxiv

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

📄 PDF Abstract BibTeX arXiv:2608.23564

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Refactoring Codebases through Library Design

2025-05-26 · Ziga Kovacic, Celine Lee, Justin Chiu, Wenting Zhao 외

Maintainable and general software allows developers to build robust applications efficiently, yet achieving these qualities often requires refactoring specialized solutions into reusable components. This challenge become…

CodeTaste: Can LLMs Generate Human-Level Code Refactorings?

2026-03-04 · Alex Thillen, Niels Mündler, Veselin Raychev, Martin Vechev arxiv

LLM coding agents can generate working code, but their solutions often accumulate complexity, duplication, and architectural debt. Human developers address such issues through refactoring: behavior-preserving program tra…

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

2026-05-08 · Mohit Raghavendra, Soham Dan, Miguel Romero Calvo, Yannis Yiming He 외 arxiv

We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test Writing (90 tasks), and Refactoring (70 tasks). SWE Atlas differs fro…

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

2026-08-10 · Yuling Shi, Jinghan Xu, Kelin Fu, Wenhao Zeng 외 hf

As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found tha…

RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

2025-03-10 · Dhruv Gautam, Spandan Garg, Jinu Jang, Neel Sundaresan 외

Recent advances in language model (LM) agents and function calling have enabled autonomous, feedback-driven systems to solve problems across various digital domains. To better understand the unique limitations of LM agen…

Specificity