paper-with-me

홈 › Papers

Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies

2025-01-24 · Lingwei Zhu, Han Wang, Yukie Nagai

Sparse continuous policies are distributions that can choose some actions at random yet keep strictly zero probability for the other actions, which are radically different from the Gaussian. They have important real-world implications, e.g. in modeling safety-critical tasks like medicine. The combination of offline reinforcement learning and sparse policies provides a novel paradigm that enables learning completely from logged datasets a safety-aware sparse policy. However, sparse policies can cause difficulty with the existing offline algorithms which require evaluating actions that fall outside of the current support. In this paper, we propose the first offline policy optimization algorithm that tackles this challenge: Fat-to-Thin Policy Optimization (FtTPO). Specifically, we maintain a fat (heavy-tailed) proposal policy that effectively learns from the dataset and injects knowledge to a thin (sparse) policy, which is responsible for interacting with the environment. We instantiate FtTPO with the general $q$-Gaussian family that encompasses both heavy-tailed and sparse policies and verify that it performs favorably in a safety-critical treatment simulation and the standard MuJoCo suite. Our code is available at \url{https://github.com/lingweizhu/fat2thin}.

📄 PDF Abstract BibTeX arXiv:2501.14373

Code (1)

lingweizhu/fat2thin 공식 구현 pytorch

Tasks

MuJoCoOffline RL

Similar Papers 제목 키워드 기반

Generative OOD-regularized Model-based Policy Optimization

2026-05-23 · Aysin Tumay, Jiahe Huang, Elise Jortberg, Rose Yu arxiv

We study sequential decision-making with offline reinforcement learning (RL). Traditional offline RL policies may result in out-of-distribution (OOD) actions when training relies only on sparse offline representations. T…

Reinforcement LearningDensity EstimationOffline RL

Offline Hierarchical Reinforcement Learning via Inverse Optimization

2024-10-10 · Carolin Schmidt, Daniele Gammelli, James Harrison, Marco Pavone 외

Hierarchical policies enable strong performance in many sequential decision-making problems, such as those with high-dimensional action spaces, those requiring long-horizon planning, and settings with sparse rewards. How…

Decision MakingHierarchical Reinforcement Learningreinforcement-learningReinforcement Learning+2

Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning

2024-05-29 · Tianle Zhang, Jiayi Guan, Lin Zhao, Yihang Li 외

Offline reinforcement learning (RL) aims to learn optimal policies from previously collected datasets. Recently, due to their powerful representational capabilities, diffusion models have shown significant potential as p…

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Trajectory-Oriented Policy Optimization with Sparse Rewards

2024-01-04 · GuoJian Wang, Faguo Wu, Xiao Zhang

Mastering deep reinforcement learning (DRL) proves challenging in tasks featuring scant rewards. These limited rewards merely signify whether the task is partially or entirely accomplished, necessitating various explorat…

continuous-controlContinuous ControlDeep Reinforcement Learning

Skill-Critic: Refining Learned Skills for Hierarchical Reinforcement Learning

2023-06-14 · Ce Hao, Catherine Weaver, Chen Tang, Kenta Kawamoto 외

Hierarchical reinforcement learning (RL) can accelerate long-horizon decision-making by temporally abstracting a policy into multiple levels. Promising results in sparse reward environments have been seen with skills, i.…

Autonomous RacingDecision MakingHierarchical Reinforcement Learningreinforcement-learning+2