paper-with-me

Papers

PEARL: Plan Exploration and Adaptive Reinforcement Learning for Multihop Tool Use

2026-01-28 · Qihao Wang, Mingzhe Lu, Jiayue Wu, Yue Hu, Yanbing Liu arxiv

Large Language Models show great potential with external tools, but face significant challenges in complex, multi-turn tool invocation. They often exhibit weak planning, tool hallucination, erroneous parameter generation, and struggle with robust interaction. To tackle these issues, we present PEARL, a novel framework to enhance LLM planning and execution for sophisticated tool use. PEARL adopts a two-stage approach: an offline phase where the agent explores tools to learn valid usage patterns and failure conditions, and an online reinforcement learning phase. In the online phase, a dedicated Planner is trained via group Relative Policy Optimization (GRPO) with a carefully designed reward function that provides distinct signals for planning quality. Experiments on the ToolHop and T-Eval benchmarks show PEARL significantly outperforms existing methods, achieving a new state-of-the-art success rate of \textbf{56.5\%} on ToolHop while maintaining a low invocation error rate. Our work marks a key advance in addressing the complex planning challenges of tool use, contributing to the development of more robust and reliable LLM-based agents.

📄 PDF Abstract BibTeX arXiv:2601.20439

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Pearl: A Production-ready Reinforcement Learning Agent

2023-12-06 · Zheqing Zhu, Rodrigo de Salvo Braz, Jalaj Bhandari, Daniel Jiang 외

Reinforcement learning (RL) is a versatile framework for optimizing long-term goals. Although many real-world problems can be formalized with RL, learning and deploying a performant RL policy requires a system designed t…

Benchmarkingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

PEARL: Preconditioner Enhancement through Actor-critic Reinforcement Learning

2025-01-18 · David Millard, Arielle Carr, Stéphane Gaudreault, Ali Baheri

We present PEARL (Preconditioner Enhancement through Actor-critic Reinforcement Learning), a novel approach to learning matrix preconditioners. Existing preconditioners such as Jacobi, Incomplete LU, and Algebraic Multig…

reinforcement-learningReinforcement Learning

Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

2025-09-19 · Zhiyu Mou, Yiqin Lv, Miao Xu, Qi Wang 외 arxiv

Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidding (AIGB), which learns a conditional generative planner from offline data, achi…

Reinforcement Learning

Low Emission Building Control with Zero-Shot Reinforcement Learning

2022-06-28 · Scott R. Jeen, Alessandro Abate, Jonathan M. Cullen

Heating and cooling systems in buildings account for 31% of global energy use, much of which are regulated by Rule Based Controllers (RBCs) that neither maximise energy efficiency nor minimise emissions by interacting op…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Low Emission Building Control with Zero-Shot Reinforcement Learning

2022-08-12 · Scott R. Jeen, Alessandro Abate, Jonathan M. Cullen

Heating and cooling systems in buildings account for 31\% of global energy use, much of which are regulated by Rule Based Controllers (RBCs) that neither maximise energy efficiency nor minimise emissions by interacting o…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)