paper-with-me

홈 › Papers

The Rank and Gradient Lost in Non-stationarity: Sample Weight Decay for Mitigating Plasticity Loss in Reinforcement Learning

2026-04-02 · Zihao Wu, Hongyao Tang, Yi Ma, Jiashun Liu, Yan Zheng, Jianye Hao arxiv

Deep reinforcement learning (RL) suffers from plasticity loss severely due to the nature of non-stationarity, which impairs the ability to adapt to new data and learn continually. Unfortunately, our understanding of how plasticity loss arises, dissipates, and can be dissolved remains limited to empirical findings, leaving the theoretical end underexplored.To address this gap, we study the plasticity loss problem from the theoretical perspective of network optimization. By formally characterizing the two culprit factors in online RL process: the non-stationarity of data distributions and the non-stationarity of targets induced by bootstrapping, our theory attributes the loss of plasticity to two mechanisms: the rank collapse of the Neural Tangent Kernel (NTK) Gram matrix and the $Θ(\frac{1}{k})$ decay of gradient magnitude. The first mechanism echoes prior empirical findings from the theoretical perspective and sheds light on the effects of existing methods, e.g., network reset, neuron recycle, and noise injection. Against this backdrop, we focus primarily on the second mechanism and aim to alleviate plasticity loss by addressing the gradient attenuation issue, which is orthogonal to existing methods. We propose Sample Weight Decay -- a lightweight method to restore gradient magnitude, as a general remedy to plasticity loss for deep RL methods based on experience replay. In experiments, we evaluate the efficacy of \methodName upon TD3, \myadded{Double DQN} and SAC with SimBa architecture in MuJoCo, \myadded{ALE} and DeepMind Control Suite tasks. The results demonstrate that \methodName effectively alleviates plasticity loss and consistently improves learning performance across various configurations of deep RL algorithms, UTD, network architectures, and environments, achieving SOTA performance on challenging DMC Humanoid tasks.

📄 PDF Abstract BibTeX arXiv:2604.01913

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Two-Channel Passive Detection Exploiting Cyclostationarity

2019-06-17 · Stefanie Horstmann, David Ramírez, Peter J. Schreier

This paper addresses a two-channel passive detection problem exploiting cyclostationarity. Given a reference channel (RC) and a surveillance channel (SC), the goal is to detect a target echo present at the surveillance a…

Vocal Bursts Valence Prediction

Assessing Cyclostationary Malware Detection via Feature Selection and Classification

2023-08-29 · Mike Nkongolo

Cyclostationarity involves periodic statistical variations in signals and processes, commonly used in signal analysis and network security. In the context of attacks, cyclostationarity helps detect malicious behaviors wi…

Anomaly DetectionClassificationfeature selectionIntrusion Detection+2

Training Non-Differentiable Networks via Optimal Transport

2026-05-03 · An T. Le arxiv

We optimize losses that jump: spiking thresholds, quantized layers, and discrete routing put jumps in the forward pass, where backpropagation does not apply. Finite differences fail: at a derivative-estimating radius, 99…

Identification of GSM and LTE Signals Using Their Second-order Cyclostationarity

2018-03-12

Automatic signal identification (ASI) has various millitary and commercial applications, such as spectrum surveillance and cognitive radio. In this paper, a novel ASI algorithm is proposed for the identification of GSM a…

Stochastic Analysis of the Diffusion Least Mean Square and Normalized Least Mean Square Algorithms for Cyclostationary White Gaussian and Non-Gaussian Inputs

2021-08-05 · Eweda Eweda, Neil J. Bershad, Jose C. M. Bermudez

The diffusion least mean square (DLMS) and the diffusion normalized least mean square (DNLMS) algorithms are analyzed for a network having a fusion center. This structure reduces the dimensionality of the resulting stoch…