paper-with-me

홈 › Papers

Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control

2023-09-26 · NeurIPS 2023 11 · Nate Rahn, Pierluca D'Oro, Harley Wiltzer, Pierre-Luc Bacon, Marc G. Bellemare

Deep reinforcement learning agents for continuous control are known to exhibit significant instability in their performance over time. In this work, we provide a fresh perspective on these behaviors by studying the return landscape: the mapping between a policy and a return. We find that popular algorithms traverse noisy neighborhoods of this landscape, in which a single update to the policy parameters leads to a wide range of returns. By taking a distributional view of these returns, we map the landscape, characterizing failure-prone regions of policy space and revealing a hidden dimension of policy quality. We show that the landscape exhibits surprising structure by finding simple paths in parameter space which improve the stability of a policy. To conclude, we develop a distribution-aware procedure which finds such paths, navigating away from noisy neighborhoods in order to improve the robustness of a policy. Taken together, our results provide new insight into the optimization, evaluation, and design of agents.

📄 PDF Abstract BibTeX arXiv:2309.14597

Code (1)

nathanrahn/return-landscapes 공식 구현 jax

Tasks

continuous-controlContinuous ControlDeep Reinforcement Learning

Similar Papers 제목 키워드 기반

Moments Matter:Stabilizing Policy Optimization using Return Distributions

2026-01-05 · Dennis Jabs, Aditya Mohan, Marius Lindauer arxiv

Deep Reinforcement Learning (RL) agents often learn policies that achieve the same episodic return yet behave very differently, due to a combination of environmental (random transitions, initial conditions, reward noise)…

Reinforcement LearningContinuous Control

Improving CMA-ES Convergence Speed, Efficiency, and Reliability in Noisy Robot Optimization Problems

2026-01-14 · Russell M. Martin, Steven H. Collins arxiv

Experimental robot optimization often requires evaluating each candidate policy for seconds to minutes. The chosen evaluation time influences optimization because of a speed-accuracy tradeoff: shorter evaluations enable …

Efficient Hill-Climber for Multi-Objective Pseudo-Boolean Optimization

2016-01-27 · Francisco Chicano, Darrell Whitley, Renato Tinos

Local search algorithms and iterated local search algorithms are a basic technique. Local search can be a stand along search methods, but it can also be hybridized with evolutionary algorithms. Recently, it has been show…

Evolutionary Algorithms

Fractal Landscapes in Policy Optimization

2023-10-24 · NeurIPS 2023 11

Policy gradient lies at the core of deep reinforcement learning (RL) in continuous domains. Despite much success, it is often observed in practice that RL training with policy gradient can fail for many reasons, even on …

Deep Reinforcement LearningReinforcement Learning (RL)

Self-Improvement Imitation with Biologically Guided Search for Protein Design Under Oracle Budgets

2026-05-26 · Ashima Khanna, Dominik Grimm arxiv

Protein sequence optimization under tight oracle budgets requires methods that explore vast combinatorial spaces while making each evaluation informative. Existing reinforcement learning and off-policy generative approac…

Reinforcement LearningProtein Design