paper-with-me

홈 › Papers

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

2026-07-27 · Hyundoo Park, Byungho Choi arxiv

Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.

📄 PDF Abstract BibTeX arXiv:2607.25152

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning a Shield from Catastrophic Action Effects: Never Repeat the Same Mistake

2022-02-19 · Shahaf S. Shperberg, Bo Liu, Peter Stone

Agents that operate in an unknown environment are bound to make mistakes while learning, including, at least occasionally, some that lead to catastrophic consequences. When humans make catastrophic mistakes, they are exp…

Continual LearningSafe Reinforcement Learning

LAGEA: Language Guided Embodied Agents for Robotic Manipulation

2025-09-27 · Abdul Monaf Chowdhury, Akm Moshiur Rahman Mazumder, Rabeya Akter, Safaeid Hossain Arib arxiv

Robotic manipulation benefits from foundation models that describe goals, but today's agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-r…

Reinforcement Learning

Stagnation Detection with Randomized Local Search

2021-01-28 · Amirhossein Rajabi, Carsten Witt

Recently a mechanism called stagnation detection was proposed that automatically adjusts the mutation rate of evolutionary algorithms when they encounter local optima. The so-called $SD-(1+1)EA$ introduced by Rajabi and …

Evolutionary Algorithms

Signals: Trajectory Sampling and Triage for Agentic Interactions

2026-04-01 · Shuguang Chen, Adil Hafeez, Salman Paracha arxiv

Agentic applications based on large language models increasingly rely on multi-step interaction loops involving planning, action execution, and environment feedback. While such systems are now deployed at scale, improvin…

Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias

2024-03-12 · Sierra Wyllie, Ilia Shumailov, Nicolas Papernot

Model-induced distribution shifts (MIDS) occur as previous model outputs pollute new model training sets over generations of models. This is known as model collapse in the case of generative models, and performative pred…

Fairness