paper-with-me

홈 › Papers

RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades

2026-05-15 · Xinbo Xu, Ruihan Yang, Haiyang Shen, Wendong Xu, Bofei Gao, Ruoyu Wu, Kean Shi, Weichu Xie, Xuanzhong Chen, Ming Wu, Jason Zeng, Michael Heinrich, Elvis Zhang, Liang Chen, Kuan Li, Baobao Chang arxiv

Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing benchmarks focus predominantly on single-issue bug fixes from Python repositories, with coarse pass/fail evaluation outcomes, and thus fail to capture long-horizon, multi-target development at real engineering scale. To address this gap, we present RoadmapBench, a benchmark of 115 long-horizon coding tasks grounded in real open-source version upgrades across 17 repositories and 5 programming languages. Each task places the agent on a source-version code snapshot and provides a multi-target roadmap instruction requiring it to implement the functionality introduced in the target version, with a median modification of 3,700 lines across 51 files. We conduct a systematic evaluation on thirteen frontier models and find that even the strongest, Claude-Opus-4.7, resolves only 39.1% of tasks, while the weakest achieves merely 5.2%, in stark contrast to existing bug-fix benchmarks, suggesting that long-horizon software development remains a largely unsolved problem.

📄 PDF Abstract BibTeX arXiv:2605.15846

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agentic Software Issue Resolution with Large Language Models: A Survey

2025-12-24 · Zhonghao Jiang, David Lo, Zhongxin Liu arxiv

Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users, and represents a key aspect of software maintenance. With the rapid developme…

Reinforcement LearningDecision Making

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

2026-05-14 · Suyoung Bae, Jaehoon Lee, Changkyu Choi, YunSeok Choi 외 arxiv

Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding agents rely on to navigate large codebases. Existing repository-level ap…

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

2026-06-09 · Liya Zhu, Jingzhe Ding, Jian Zhang, Jianbo Xue 외 arxiv

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to co…

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

2026-02-26 · Yujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan 외 arxiv

Large Language Models (LLMs) are increasingly used as autonomous agents in complex, long-horizon applications, where effective memory is critical for sustained performance. Yet existing memory benchmarks are largely dial…

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

2026-05-28 · Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding 외 arxiv

Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We …