ADARL: Adaptive Low-Rank Structures for Robust Policy Learning under Uncertainty
Robust reinforcement learning (Robust RL) seeks to handle epistemic uncertainty in environment dynamics, but existing approaches often rely on nested min--max optimization, which is computationally expensive and yields overly conservative policies. We propose \textbf{Adaptive Rank Representation (AdaRL)}, a bi-level optimization framework that improves robustness by aligning policy complexity with the intrinsic dimension of the task. At the lower level, AdaRL performs policy optimization under fixed-rank constraints with dynamics sampled from a Wasserstein ball around a centroid model. At the upper level, it adaptively adjusts the rank to balance the bias--variance trade-off, projecting policy parameters onto a low-rank manifold. This design avoids solving adversarial worst-case dynamics while ensuring robustness without over-parameterization. Empirical results on MuJoCo continuous control benchmarks demonstrate that AdaRL not only consistently outperforms fixed-rank baselines (e.g., SAC) and state-of-the-art robust RL methods (e.g., RNAC, Parseval), but also converges toward the intrinsic rank of the underlying tasks. These results highlight that adaptive low-rank policy representations provide an efficient and principled alternative for robust RL under model uncertainty.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningContinuous ControlSimilar Papers 제목 키워드 기반
AdaRL: What, Where, and How to Adapt in Transfer Reinforcement Learning
One practical challenge in reinforcement learning (RL) is how to make quick adaptations when faced with new environments. In this paper, we propose a principled framework for adaptive RL, called \textit{AdaRL}, that adap…
Atari Gamesreinforcement-learningReinforcement Learning (RL)Transfer Reinforcement LearningRadarly : \'ecouter et analyser le web conversationnel en temps r\'eel (Real time listening and analysis of the social web using Radarly)
De par le contexte conversationnel digital, l{'}outil Radarly a {\'e}t{\'e} con{\c{c}}u pour permettre de traiter de grands volumes de donn{\'e}es h{\'e}t{\'e}rog{\`e}nes en temps r{\'e}el, de g{\'e}n{\'e}rer de nouveaux…
RadarLCD: Learnable Radar-based Loop Closure Detection Pipeline
Loop Closure Detection (LCD) is an essential task in robotics and computer vision, serving as a fundamental component for various applications across diverse domains. These applications encompass object recognition, imag…
Image RetrievalLoop Closure DetectionObject RecognitionRadar odometryRadarLoc: Learning to Relocalize in FMCW Radar
Relocalization is a fundamental task in the field of robotics and computer vision. There is considerable work in the field of deep camera relocalization, which directly estimates poses from raw images. However, learning-…
Camera RelocalizationOptimizing adaptive sampling via Policy Ranking
Efficient sampling in biomolecular simulations is critical for accurately capturing the complex dynamical behaviors of biological systems. Adaptive sampling techniques aim to improve efficiency by focusing computational …