paper-with-me

홈 › Papers

PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents

2026-06-15 · Zhenbang Du, Jun Luo, Zhiwei Zheng, Xiangchi Yuan, Kejing Xia, Dachuan Shi, Qirui Jin, Qijia He, Shaofeng Zou, Yingbin Liang, Wenke Lee arxiv

Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents is challenging, as reinforcement learning often suffers from sparse rewards and weak credit assignment despite matching the prompt-only inference setting, while supervised fine-tuning on expert traces provides dense process supervision but can over-constrain the model to fixed trajectories. To tackle this, we propose PACT, a Privileged trAce Co-Training framework for multi-turn tool-use agents. The key idea is to use expert traces only as training-time optimization signals rather than rollout-time hints. PACT keeps rollout generation prompt-only, then uses expert traces to guide optimization through two complementary signals: a trace-conditioned RL surrogate that evaluates prompt-only rollouts under expert-trace context, and a component-aware SFT loss that supervises reasoning prefixes and tool-calls with annealed strength. To reduce over-reliance on the training-only trace context, PACT further introduces a prompt-only anchoring. We also provide a latent-trace view that connects the two trace-based objectives and explains how expert traces can guide optimization without being used during rollout generation. Experiments on FTRL, BFCL, and ToolHop show that PACT consistently improves over strong SFT- and RL-based baselines, highlighting the value of privileged trace co-training for multi-turn tool-use learning.

📄 PDF Abstract BibTeX arXiv:2606.16215

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

2026-04-12 · Hao Wang, Guozhi Wang, Han Xiao, Yufeng Zhou 외 arxiv

Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited by sparse rewards and long horizons. On-policy self-distillation (OPSD)…

Reinforcement Learning

HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

2026-06-10 · Haoran Liu, Yuwei Zhang, Xiyao Li, Bohan Lyu 외 arxiv

Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to determine credit assignments for each intermediate turns. Recent on-po…

Reinforcement Learning

Vision-Based Deep Reinforcement Learning of UAV Autonomous Navigation Using Privileged Information

2024-12-09 · Junqiao Wang, Zhongliang Yu, Dong Zhou, Jiaqi Shi 외

The capability of UAVs for efficient autonomous navigation and obstacle avoidance in complex and unknown environments is critical for applications in agricultural irrigation, disaster relief and logistics. In this paper,…

Autonomous NavigationBenchmarkingDeep Reinforcement Learningreinforcement-learning+1

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

2026-08-05 · Junzhuo Liu, Weiwei Li, Jun Ling, Peng Wang hf

Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as succ…

TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment

2026-05-11 · Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen, Yulan Hu 외 arxiv

On-policy self-distillation (self-OPD) densifies reinforcement learning with verifiable rewards (RLVR) by letting a policy teach itself under privileged context. We find that when this guidance spans the full response, a…

Reinforcement Learning