paper-with-me

홈 › Papers

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

2026-01-22 · Zhiwei Zhang, Fei Zhao, Rui Wang, Zezhong Wang, Bin Liang, Jiakang Wang, Yao Hu, Shaosheng Cao, Kam-Fai Wong arxiv

Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models often fall into repetitive invalid re-invocations instead of interpreting the feedback and recovering. This failure mode persists because current training paradigms do not explicitly teach models how to recover from execution errors. In particular, standard reinforcement learning (RL) collapses rich failure experience into sparse negative rewards, while pre-collected error-correction datasets become mismatched to the policy's evolving failure modes. To bridge this gap, we propose Fission-GRPO, a framework that converts execution errors into on-policy corrective supervision within the RL training loop. Our core mechanism fissions each failed trajectory into a new training instance by augmenting it with diagnostic feedback from a fine-tuned Error Simulator, then resampling multiple recovery rollouts on-policy. This enables the model to learn from the precise errors it makes during exploration, rather than from static, pre-collected error cases. On BFCL v4 Multi-Turn, Fission-GRPO improves the error recovery rate of Qwen3-8B by 5.7% absolute and overall accuracy by 4.0% (from 42.75% to 46.75%), outperforming both RL baselines and specialized tool-use agents. The method further generalizes to TAU-Bench and TAU2-Bench, achieving leading results across most settings with gains up to +17.4%.

📄 PDF Abstract BibTeX arXiv:2601.15625

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

2026-05-18 · Li Wang, Xiaohan Wang, Xiaodong Lu, Zipeng Zhang 외 arxiv

Large language models (LLMs) have increasingly leveraged tool invocation to enhance their reasoning capabilities. However, existing approaches typically tightly couple tool invocation with immediate execution. Such immed…

Mathematical Reasoning

Data fission: splitting a single data point

2021-12-21 · James Leiner, Boyan Duan, Larry Wasserman, Aaditya Ramdas

Suppose we observe a random vector $X$ from some distribution $P$ in a known family with unknown parameters. We ask the following question: when is it possible to split $X$ into two parts $f(X)$ and $g(X)$ such that neit…

Additive modelsBayesian Inference

MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery

2026-05-07 · Nanjie Yao, Junlong Ren, Wenhao Shen, Hao Wang arxiv

This paper studies full-body 3D human motion recovery from head-mounted device signals. Existing diffusion-based methods often rely on global distribution matching, leading to local joint reconstruction errors. We propos…

Reinforcement Learning

WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning

2026-01-06 · Xinmiao Yu, Liwen Zhang, Xiaocheng Feng, Yong Jiang 외 arxiv

Large Language Model(LLM)-based agents have shown strong capabilities in web information seeking, with reinforcement learning (RL) becoming a key optimization paradigm. However, planning remains a bottleneck, as existing…

Reinforcement Learning

Learning to Refine: An Agentic RL Approach for Iterative SPARQL Query Construction

2025-11-14 · Floris Vossebeld, Shenghui Wang arxiv

Generating complex, logically-sound SPARQL queries for multi-hop questions remains a critical bottleneck for Knowledge Graph Question Answering, as the brittle nature of one-shot generation by Large Language Models (LLMs…

Graph Question AnsweringReinforcement LearningKnowledge Graphs