paper-with-me

홈 › Papers

Probe-and-Refine Tuning of Repository Guidance for Coding Agents

2026-06-18 · Asa Shepard, Jeannie Albrecht arxiv

LLM-based coding agents need higher-level operational knowledge about a repository (which files house which subsystems, how to run the test suite, which workflows have historically led to wrong fixes) that does not exist in the code itself. Engineers typically maintain AGENTS.md files to supply this context as instructions for coding agents, but whether they help is contested: recent studies disagree on whether LLM-generated guidance improves or harms agent performance. In this paper we show that how the guidance is produced is the decisive variable, and introduce probe-and-refine tuning: a procedure that uses synthetic bug-fix probes to iteratively diagnose and patch a repository's guidance file through single-shot LLM calls, with no agent loop or tool use during tuning. On SWE-bench Verified across four independent trials with Qwen3.5-35B-A3B at 200 steps, probe-and-refine achieves 33.0% mean resolve rate vs. 28.3% for the static knowledge base used to initialize it and 25.5% for an unguided baseline (p < 0.001 for both probe-and-refine contrasts). The improvement comes from coverage rather than precision: refined guidance produces evaluable patches for 14.5 percentage points (pp) more instances while per-patch precision remains statistically constant (~59%, p = 0.119), showing that improved guidance helps agents reach the correct file rather than improving the quality of the changes they make. Further, a step-budget experiment shows that guidance is what lets the agent use a larger step budget productively, and a cross-model experiment with NVIDIA-Nemotron-3-Nano-30B-A3B finds that the tuning loop degrades when the model cannot generate sufficiently diagnostic output, though per-patch precision remains constant even then.

📄 PDF Abstract BibTeX arXiv:2606.20512

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents

2026-05-09 · Chenyu Zhao, Shenglin Zhang, Yihang Lin, Wenwei Gu 외 arxiv

Software engineering agents are increasingly deployed in evaluable engineering environments, yet post-failure recovery remains costly, manual, and ad hoc. Existing systems expose traces or generate follow-up feedback, bu…

EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance

2025-04-17 · CVPR 2025 1 · Yang Yue, Yulin Wang, Haojun Jiang, Pan Liu 외

Echocardiography is crucial for cardiovascular disease detection but relies heavily on experienced sonographers. Echocardiography probe guidance systems, which provide real-time movement instructions for acquiring standa…

Anatomy

Needle in the Repo: A Benchmark for Maintainability in AI-Generated Repository Edits

2026-03-29 · Haichao Zhu, Qian Zhang, Jiyuan Wang, Zhaorui Yang 외 arxiv

AI coding agents can now complete complex programming tasks, but existing evaluations largely emphasize behavioral correctness and often overlook maintainability risks such as weak modularity or testability. We present N…

Code Generation

RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph

2024-10-03 · Siru Ouyang, Wenhao Yu, Kaixin Ma, Zilin Xiao 외

Large Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks. Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic …

Code Generation

AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback

2026-05-20 · Miaobo Hu, Shuhao Hu, Bokun Wang, Ruohan Wang 외 arxiv

Reinforcement learning improves LLM reasoning, but PPO/GRPO typically use fixed clipping and decoding temperature, which makes training brittle and tuning-heavy. We propose Adaptive Group Policy Optimization (AGPO), a cr…

Reinforcement Learning