paper-with-me

홈 › Papers

ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing

2025-11-18 · Ankush Kadu, Aswanth Krishnan arxiv

We present ReflexGrad, a dual-process architecture for within-episode failure recovery in LLM agents without demonstrations. When agents commit to a wrong approach early and exhaust the step budget, the post-failure trajectory contains the information to escape -- but no published architecture acts on it within a single episode. ReflexGrad routes between a fast process (TextGrad-style continuous refinement every $k{=}3$ steps) and a slow process (Reflexion-style causal diagnosis when $m{=}5$ consecutive low-progress scores fire a routing gate). A deterministic priority merge keeps the natural-language policy coherent, and each slow activation emits three observable artifacts: a reproducible trigger, a causal diagnostic, and a verified fix. On ALFWorld 134 tasks, $n{=}10$ seeds, no demonstrations, ReflexGrad lifts Qwen-3-8B from $35.1\%$ to $75.4\%$ ($+40.3$pp), beating compute-matched 1-shot LATS by $+2.7$pp ($p{\approx}0.01$), ToT by $+5.7$pp ($p{<}10^{-4}$), and Self-Refine by $+6.7$pp ($p{<}10^{-5}$); on GPT-5 the lift is $46.3{\to}88.1\%$ ($+41.8$pp). The $1.5$pp cross-model difference is within seed noise ($p{\approx}0.13$), suggesting that the routing mechanism, rather than model scale, is the primary source of the gain. Code, prompts, per-seed logs, and sensitivity sweeps are released.

📄 PDF Abstract BibTeX arXiv:2511.14584

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Failure Modes of Deep Multi-Agent RL in Asynchronous Pricing: Reproducible Triggers, Trace Diagnostics, and a Partial Fix

2026-06-03 · Shree Murthy, Rohan Pandey arxiv

We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation between competing DDPG agents, and (ii) actor--critic instability at high …

Multi-agent Reinforcement Learning

Living-Harness Is an Interactive-Agent Evolver

2026-07-29 · Yuetian Du, Yucheng Wang, He Xu, Jiexu Xu 외 arxiv

Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness…

Conditional Multi-Stage Failure Recovery for Embodied Agents

2025-07-08 · Youmna Farag, Svetlana Stoyanchev, Mohan Li, Simon Keizer 외

Embodied agents performing complex tasks are susceptible to execution failures, motivating the need for effective failure recovery mechanisms. In this work, we introduce a conditional multistage failure recovery framewor…

CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery

2026-07-03 · Safia Fatima, Kai Olav Ellefsen, Leon Moonen arxiv

Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenarios. RL policies often stall in failure states, spending up to 70% of an…

Reinforcement Learning

Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents

2026-05-09 · Chenyu Zhao, Shenglin Zhang, Yihang Lin, Wenwei Gu 외 arxiv

Software engineering agents are increasingly deployed in evaluable engineering environments, yet post-failure recovery remains costly, manual, and ad hoc. Existing systems expose traces or generate follow-up feedback, bu…