paper-with-me

홈 › Papers

Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents

2026-04-29 · Anh Ta, Junjie Zhu, Shahin Shayandeh arxiv

Tool-calling agents are evaluated on tool selection, parameter accuracy, and scope recognition, yet LLM trajectory assessments remain inherently post-hoc. Disconnected from the active execution loop, such assessments identify errors that are usually addressed through prompt-tuning or retraining, and fundamentally cannot course-correct the agent in real time. To close this gap, we move evaluation into the execution loop at inference time: a specialized reviewer agent evaluates provisional tool calls prior to execution, shifting the paradigm from post-hoc recovery to proactive evaluation and error mitigation. In practice, this architecture establishes a clear separation of concerns between the primary execution agent and a secondary review agent. As with any multi-agent system, the reviewer can introduce new errors while correcting others, yet no prior work to our knowledge has systematically measured this tradeoff. To quantify this tradeoff, we introduce Helpfulness-Harmfulness metrics: helpfulness measures the percentage of base agent errors that feedback corrects; harmfulness measures the percentage of correct responses that feedback degrades. These metrics directly inform reviewer design by revealing whether a given model or prompt provides net positive value. We evaluate our approach on BFCL (single-turn) and Tau2-Bench (multi-turn stateful scenarios), achieving +5.5% on irrelevance detection and +7.1% on multi-turn tasks. Our metrics reveal that reviewer model choice is critical: the reasoning model o3-mini achieves a 3:1 benefit-to-risk ratio versus 2.1:1 for GPT-4o. Automated prompt optimization via GEPA provides an additional +1.5-2.8%. Together, these results demonstrate a core advantage of separating execution and review: the reviewer can be systematically improved through model selection and prompt optimization, without retraining the base agent.

📄 PDF Abstract BibTeX arXiv:2604.27233

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Flexible Multi-Agent LLM-Human Framework for Fast Human Validated Tool Building

2025-12-01 · Daull Xavier, Patrice Bellot, Emmanuel Bruno, Vincent Martin 외 arxiv

We introduce CollabToolBuilder, a flexible multiagent LLM framework with expert-in-the-loop (HITL) guidance that iteratively learns to create tools for a target goal, aligning with human intent and process, while minimiz…

Domain Adaptation

Re-ReST: Reflection-Reinforced Self-Training for Language Agents

2024-06-03 · Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu, Kai-Wei Chang 외

Finetuning language agents with reasoning-action trajectories is effective, but obtaining these trajectories from human annotations or stronger models is costly and sometimes impractical. In this paper, we investigate th…

Code GenerationImage GenerationMulti-hop Question AnsweringQuestion Answering+4

Experiential Reinforcement Learning

2026-02-15 · Taiwei Shi, Sihao Chen, Bowen Jiang, Linxin Song 외 arxiv

Reinforcement learning has become the central approach for language models (LMs) to learn from environmental reward or feedback. In practice, the environmental feedback is usually sparse and delayed. Learning from such s…

Reinforcement Learning

Distilling Feedback into Memory-as-a-Tool

2026-01-09 · Víctor Gallego arxiv

We propose a framework that amortizes the cost of inference-time reasoning by converting transient critiques into retrievable guidelines, through a file-based memory system and agent-controlled tool calls. We evaluate th…

FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing

2026-02-02 · Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Qika Lin 외 arxiv

In recent years, agentic workflows have been widely applied to solve complex human tasks. However, existing workflow construction still faces key challenges, including human-dependent workflow construction, the lack of g…

Reinforcement Learning