paper-with-me

홈 › Papers

Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?

2025-10-01 · Haizhong Zheng, Jiawei Zhao, Beidi Chen arxiv

Reinforcement learning has been central to recent advances in large language model reasoning, but most algorithms rely on on-policy training that demands fresh rollouts at every update, limiting efficiency and scalability. Asynchronous RL systems alleviate this by decoupling rollout generation from training, yet their effectiveness hinges on tolerating large staleness in rollout data, a setting where existing methods either degrade in performance or collapse. We revisit this challenge and uncover a prosperity-before-collapse phenomenon: stale data can be as informative as on-policy data if exploited properly. Building on this insight, we introduce M2PO (Second-Moment Trust Policy Optimization), which constrains the second moment of importance weights to suppress only extreme outliers while preserving informative updates. Notably, M2PO sharply reduces the fraction of clipped tokens under high staleness (from 1.22% to 0.06% over training), precisely masking high-variance tokens while maintaining stable optimization. Extensive evaluation across six models (from 1.7B to 32B) and eight benchmarks shows that M2PO delivers stable off-policy training even with data stale by at least 256 model updates and matches on-policy performance.

📄 PDF Abstract BibTeX arXiv:2510.01161

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

2026-07-01 · Jingwei Song, Haofeng Xu, Jie Xiao, Chengke Bao 외 arxiv

High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GR…

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

2026-05-29 · Resham Joshi arxiv

AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence, extracted facts, retrieved context, and answer policy into one opaqu…

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

2026-07-21 · Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li 외 arxiv

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing…

Reinforcement Learning

AsyncOPD: How Stale Can On-Policy Distillation Be?

2026-06-23 · Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjun Kang 외 arxiv

On-policy distillation (OPD) trains a student on its own rollouts guided by teacher feedback and is becoming increasingly important for large language model (LLM) post-training. Like reinforcement learning (RL), however,…

Reinforcement Learning

AgentComm-Bench: Stress-Testing Cooperative Embodied AI Under Latency, Packet Loss, and Bandwidth Collapse

2026-03-18 · Aayam Bansal, Ishaan Gangwani arxiv

Cooperative multi-agent methods for embodied AI are almost universally evaluated under idealized communication: zero latency, no packet loss, and unlimited bandwidth. Real-world deployment on robots with wireless links, …

Autonomous Vehicles