paper-with-me

Papers

DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle

2026-01-27 · Yuheng Tang, Kaijie Zhu, Bonan Ruan, Chuqi Zhang, Michael Yang, Hongwei Li, Suyue Guo, Tianneng Shi, Zekun Li, Christopher Kruegel, Giovanni Vigna, Dawn Song, William Yang Wang, Lun Wang, Yangruibo Ding, Zhenkai Liang, Wenbo Guo arxiv

Even though demonstrating extraordinary capabilities in code generation and software issue resolving, AI agents' capabilities in the full software DevOps cycle are still unknown. Different from pure code generation, handling the DevOps cycle in real-world software, including developing, deploying, and managing, requires analyzing large-scale projects, understanding dynamic program behaviors, leveraging domain-specific tools, and making sequential decisions. However, existing benchmarks focus on isolated problems and lack environments and tool interfaces for DevOps. We introduce DevOps-Gym, the first end-to-end benchmark for evaluating AI agents across core DevOps workflows: build and configuration, monitoring, issue resolving, and test generation. DevOps-Gym includes 700+ real-world tasks collected from 30+ projects in Java and Go. We develop a semi-automated data collection mechanism with rigorous and non-trivial expert efforts in ensuring the task coverage and quality. Our evaluation of state-of-the-art models and agents reveals fundamental limitations: they struggle with issue resolving and test generation in Java and Go, and remain unable to handle new tasks such as monitoring and build and configuration. These results highlight the need for essential research in automating the full DevOps cycle with AI agents.

📄 PDF Abstract BibTeX arXiv:2601.20882

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

The Role of DevOps in Enhancing Enterprise Software Delivery Success through R&D Efficiency and Source Code Management

2024-11-04 · Jun Cui

This study examines the impact of DevOps practices on enterprise software delivery success, focusing on enhancing R&D efficiency and source code management (SCM). Using a qualitative methodology, data were collected from…

Management

AI for DevSecOps: A Landscape and Future Opportunities

2024-04-07 · Michael Fu, Jirat Pasuksmit, Chakkrit Tantithamthavorn

DevOps has emerged as one of the most rapidly evolving software development paradigms. With the growing concerns surrounding security in software systems, the DevSecOps paradigm has gained prominence, urging practitioner…

On Continuous Integration / Continuous Delivery for Automated Deployment of Machine Learning Models using MLOps

2022-02-07 · Satvik Garg, Pradyumn Pundir, Geetanjali Rathee, P. K. Gupta 외

Model deployment in machine learning has emerged as an intriguing field of research in recent years. It is comparable to the procedure defined for conventional software development. Continuous Integration and Continuous …

BIG-bench Machine Learning

Developing and Operating Artificial Intelligence Models in Trustworthy Autonomous Systems

2020-03-11 · Silverio Martínez-Fernández, Xavier Franch, Andreas Jedlitschka, Marc Oriol 외

Companies dealing with Artificial Intelligence (AI) models in Autonomous Systems (AS) face several problems, such as users' lack of trust in adverse or unknown conditions, gaps between software engineering and AI model d…

Semantic Similarity-Based Clustering of Findings From Security Testing Tools

2022-11-20 · Phillip Schneider, Markus Voggenreiter, Abdullah Gulraiz, Florian Matthes

Over the last years, software development in domains with high security demands transitioned from traditional methodologies to uniting modern approaches from software development and operations (DevOps). Key principles o…

ClusteringSemantic SimilaritySemantic Textual Similarity