paper-with-me

홈 › Papers

Moments Matter:Stabilizing Policy Optimization using Return Distributions

2026-01-05 · Dennis Jabs, Aditya Mohan, Marius Lindauer arxiv

Deep Reinforcement Learning (RL) agents often learn policies that achieve the same episodic return yet behave very differently, due to a combination of environmental (random transitions, initial conditions, reward noise) and algorithmic (minibatch selection, exploration noise) factors. In continuous control tasks, even small parameter shifts can produce unstable gaits, complicating both algorithm comparison and real-world transfer. Previous work has shown that such instability arises when policy updates traverse noisy neighborhoods and that the spread of post-update return distribution $R(θ)$, obtained by repeatedly sampling minibatches, updating $θ$, and measuring final returns, is a useful indicator of this noise. Although explicitly constraining the policy to maintain a narrow $R(θ)$ can improve stability, directly estimating $R(θ)$ is computationally expensive in high-dimensional settings. We propose an alternative that takes advantage of environmental stochasticity to mitigate update-induced variability. Specifically, we model state-action return distribution through a distributional critic and then bias the advantage function of PPO using higher-order moments (skewness and kurtosis) of this distribution. By penalizing extreme tail behaviors, our method discourages policies from entering parameter regimes prone to instability. We hypothesize that in environments where post-update critic values align poorly with post-update returns, standard PPO struggles to produce a narrow $R(θ)$. In such cases, our moment-based correction narrows $R(θ)$, improving stability by up to 75% in Walker2D, while preserving comparable evaluation returns.

📄 PDF Abstract BibTeX arXiv:2601.01803

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningContinuous Control

Similar Papers 제목 키워드 기반

Adaptive Student's t-distribution with method of moments moving estimator for nonstationary time series

2023-04-06 · Jarek Duda

The real life time series are usually nonstationary, bringing a difficult question of model adaptation. Classical approaches like ARMA-ARCH assume arbitrary type of dependence. To avoid their bias, we will focus on recen…

PhilosophyTime Series

Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk

2026-01-30 · Rohan Tangri, Jan-Peter Calliess arxiv

We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems. We employ Cantelli's inequality to obtain a tractable, conservative and smooth bound on…

Reinforcement Learning

Scale matters: The daily, weekly and monthly volatility and predictability of Bitcoin, Gold, and the S&P 500

2021-02-28 · Nassim Dehouche

A reputation of high volatility accompanies the emergence of Bitcoin as a financial asset. This paper intends to nuance this reputation and clarify our understanding of Bitcoin's volatility. Using daily, weekly, and mont…

Market-Based Probability of Stock Returns

2023-02-06 · Victor Olkhov

This paper describes the dependence of market-based statistical moments of returns on statistical moments and correlations of the current and past trade values. We use Markowitz's definition of value weighted return of a…

Time SeriesTime Series Analysis

Stabilizing Policy Optimization via Logits Convexity

2026-03-01 · Hongzhan Chen, Tao Yang, Yuhua Zhu, Shiping Gao 외 arxiv

While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to supervised fine-tuning (SFT). In this work, w…

Reinforcement Learning