paper-with-me

Papers

Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization Method

2020-12-01 · NeurIPS 2020 12 · Qi Zhou, Yufei Kuang, Zherui Qiu, Houqiang Li, Jie Wang

Many recent reinforcement learning (RL) methods learn stochastic policies with entropy regularization for exploration and robustness. However, in continuous action spaces, integrating entropy regularization with expressive policies is challenging and usually requires complex inference procedures. To tackle this problem, we propose a novel regularization method that is compatible with a broad range of expressive policy architectures. An appealing feature is that, the estimation of our regularization terms is simple and efficient even when the policy distributions are unknown. We show that our approach can effectively promote the exploration in continuous action spaces. Based on our regularization, we propose an off-policy actor-critic algorithm. Experiments demonstrate that the proposed algorithm outperforms state-of-the-art regularized RL methods in continuous control tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Entropy Regularization 설명 없음

Similar Papers 제목 키워드 기반

Improving Generative Behavior Cloning via Self-Guidance and Adaptive Chunking

2025-10-14 · Junhyuk So, Chiwoong Lee, Shinyoung Lee, Jungseul Ok 외 arxiv

Generative Behavior Cloning (GBC) is a simple yet effective framework for robot learning, particularly in multi-task settings. Recent GBC methods often employ diffusion policies with open-loop (OL) control, where actions…

Methods for Sparse and Low-Rank Recovery under Simplex Constraints

2016-05-02 · Ping Li, Syama Sundar Rangapuram, Martin Slawski

The de-facto standard approach of promoting sparsity by means of $\ell_1$-regularization becomes ineffective in the presence of simplex constraints, i.e.,~the target is known to have non-negative entries summing up to a …

compressed sensingDensity EstimationPortfolio OptimizationQuantum State Tomography

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning

2025-10-10 · Zhenglin Wan, Jingxuan Wu, Xingrui Yu, Chubin Zhang 외 arxiv

Flow Matching (FM) has shown remarkable ability in modeling complex distributions and achieves strong performance in offline imitation learning for cloning expert behaviors. However, despite its behavioral cloning expres…

Reinforcement Learning

Viability of Future Actions: Robust Safety in Reinforcement Learning via Entropy Regularization

2025-06-12 · Pierre-François Massiani, Alexander von Rohr, Lukas Haverbeck, Sebastian Trimpe

Despite the many recent advances in reinforcement learning (RL), the question of learning policies that robustly satisfy state constraints under unknown disturbances remains open. In this paper, we offer a new perspectiv…

Reinforcement Learning (RL)

Increasing Entropy to Boost Policy Gradient Performance on Personalization Tasks

2023-10-09 · Andrew Starnes, Anton Dereventsov, Clayton Webster

In this effort, we consider the impact of regularization on the diversity of actions taken by policies generated from reinforcement learning agents trained using a policy gradient. Policy gradient agents are prone to ent…

Diversity