paper-with-me

홈 › Papers

AgentDevel: Reframing Self-Evolving LLM Agents as Release Engineering

2026-01-08 · Di Zhang arxiv

Recent progress in large language model (LLM) agents has largely focused on embedding self-improvement mechanisms inside the agent or searching over many concurrent variants. While these approaches can raise aggregate scores, they often yield unstable and hard-to-audit improvement trajectories, making it difficult to guarantee non-regression or to reason about failures across versions. We reframe agent improvement as \textbf{release engineering}: agents are treated as shippable artifacts, and improvement is externalized into a regression-aware release pipeline. We introduce \textbf{AgentDevel}, a release engineering pipeline that iteratively runs the current agent, produces implementation-blind, symptom-level quality signals from execution traces, synthesizes a single release candidate (RC) via executable diagnosis, and promotes it under flip-centered gating. AgentDevel features three core designs: (i) an implementation-blind LLM critic that characterizes failure appearances without accessing agent internals, (ii) script-based executable diagnosis that aggregates dominant symptom patterns and produces auditable engineering specifications, and (iii) flip-centered gating that prioritizes pass to fail regressions and fail to pass fixes as first-class evidence. Unlike population-based search or in-agent self-refinement, AgentDevel maintains a single canonical version line and emphasizes non-regression as a primary objective. Experiments on execution-heavy benchmarks demonstrate that AgentDevel yields stable improvements with significantly fewer regressions while producing reproducible, auditable artifacts. Overall, AgentDevel provides a practical development discipline for building, debugging, and releasing LLM agents as software development.

📄 PDF Abstract BibTeX arXiv:2601.04620

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

2024-02-18 · Siyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu Wei 외

This paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs), aiming for a more accurate assessment of their capabilities and limitations. We utilize a mu…

Model Selection

AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models

2025-09-30 · Yixu Wang, Xin Wang, Yang Yao, Xinyuan Li 외 arxiv

The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, existing static benchmarks are ill-equipped to address the dynamic nature of…

TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments

2026-08-14 · Qingren Yao, Yaxuan Kong, Yuqi Nie, Yichen Li 외 arxiv

Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing time series QA benchmarks mostly rely on f…

Time Series Analysis

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

2026-08-06 · Jiale Han, Xiang Li, Jing Qian, Wenyuan Gu 외 arxiv

Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through …

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

2026-07-31 · Dong Yan, Jian Liang, Dapeng Hu, Ran He 외 hf

Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-e…