paper-with-me

Papers

Delightful Distributed Policy Gradient

2026-03-20 · Ian Osband arxiv

Distributed reinforcement learning trains on data from stale, buggy, or mismatched actors, producing actions with high surprisal (negative log-probability) under the learner's policy. The core difficulty is not surprising data per se, but \emph{negative learning from surprising data}. High-surprisal failures can dominate finite-batch updates through large perpendicular components, while high-surprisal successes reveal opportunities the current policy would otherwise miss. The \textit{Delightful Policy Gradient} (DG) separates these cases by gating each update with delight, the product of advantage and surprisal, suppressing rare failures and preserving rare successes without behavior probabilities. In a tabular analysis, DG suppresses the perpendicular second moment of high-surprisal failures by a policy-overlap factor that vanishes as the learner improves. The advantage sign is essential for surprisal-based filtering: any learner-probability-only gate that suppresses rare failures also suppresses rare successes. On MNIST with simulated staleness, DG without off-policy correction outperforms importance-weighted PG with exact behavior probabilities. On a transformer sequence task with staleness, actor bugs, reward corruption, and rare discovery, DG often achieves nearly order-of-magnitude lower error. When all four frictions act simultaneously, its sample-efficiency advantage is order-of-magnitude and grows with task complexity.

📄 PDF Abstract BibTeX arXiv:2603.20521

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Delightful Policy Gradient

2026-03-15 · Ian Osband arxiv

Standard policy gradients weight each sampled action by advantage alone, regardless of how likely that action was under the current policy. This creates two pathologies: within a single decision context (e.g. one image o…

Continuous Control

Delightful Gradients Accelerate Corner Escape

2026-05-12 · Jincheng Mei, Ian Osband arxiv

Softmax policy gradient converges at $O(1/t)$, but its transient behavior near sub-optimal corners of the simplex can be exponentially slow. The bottleneck is self-trapping: negative-advantage actions reinforce the corne…

Does This Gradient Spark Joy?

2026-03-20 · Ian Osband arxiv

Policy gradient computes a backward pass for every sample, even though the backward pass is expensive and most samples carry little learning value. The Delightful Policy Gradient (DG) provides a forward-pass signal of le…

Distributed Policy Gradient with Variance Reduction in Multi-Agent Reinforcement Learning

2021-11-25 · Xiaoxiao Zhao, Jinlong Lei, Li Li, Jie Chen

This paper studies a distributed policy gradient in collaborative multi-agent reinforcement learning (MARL), where agents over a communication network aim to find the optimal policy to maximize the average of all agents'…

Multi-agent Reinforcement Learningreinforcement-learningReinforcement Learning (RL)Stochastic Optimization

Communication-Efficient Policy Gradient Methods for Distributed Reinforcement Learning

2018-12-07 · Tianyi Chen, Kaiqing Zhang, Georgios B. Giannakis, Tamer Başar

This paper deals with distributed policy optimization in reinforcement learning, which involves a central controller and a group of learners. In particular, two typical settings encountered in several applications are co…

Distributed ComputingMulti-agent Reinforcement LearningPolicy Gradient Methodsreinforcement-learning+2