paper-with-me

홈 › Papers

Maximum Entropy Reinforcement Learning with Mixture Policies

2021-03-18 · Nir Baram, Guy Tennenholtz, Shie Mannor

Mixture models are an expressive hypothesis class that can approximate a rich set of policies. However, using mixture policies in the Maximum Entropy (MaxEnt) framework is not straightforward. The entropy of a mixture model is not equal to the sum of its components, nor does it have a closed-form expression in most cases. Using such policies in MaxEnt algorithms, therefore, requires constructing a tractable approximation of the mixture entropy. In this paper, we derive a simple, low-variance mixture-entropy estimator. We show that it is closely related to the sum of marginal entropies. Equipped with our entropy estimator, we derive an algorithmic variant of Soft Actor-Critic (SAC) to the mixture policy case and evaluate it on a series of continuous control tasks.

📄 PDF Abstract BibTeX arXiv:2103.10176

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous Controlreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Maximum Entropy Diverse Exploration: Disentangling Maximum Entropy Reinforcement Learning

2019-11-03 · Andrew Cohen, Lei Yu, Xingye Qiao, Xiangrong Tong

Two hitherto disconnected threads of research, diverse exploration (DE) and maximum entropy RL have addressed a wide range of problems facing reinforcement learning algorithms via ostensibly distinct mechanisms. In this …

Diversityreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Evidence on the Regularisation Properties of Maximum-Entropy Reinforcement Learning

2025-01-28 · Rémy Hosseinkhan Boucher, Onofrio Semeraro, Lionel Mathelin

The generalisation and robustness properties of policies learnt through Maximum-Entropy Reinforcement Learning are investigated on chaotic dynamical systems with Gaussian noise on the observable. First, the robustness un…

Learning Theory

DIME:Diffusion-Based Maximum Entropy Reinforcement Learning

2025-02-04 · Onur Celik, Zechu Li, Denis Blessing, Ge Li 외

Maximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties. Traditionally, policies are parameterized using Gaussian distributions, which signif…

reinforcement-learningReinforcement Learning

Maximum Causal Tsallis Entropy Imitation Learning

2018-05-22 · NeurIPS 2018 12 · Kyungjae Lee, Sungjoon Choi, Songhwai Oh

In this paper, we propose a novel maximum causal Tsallis entropy (MCTE) framework for imitation learning which can efficiently learn a sparse multi-modal policy distribution from demonstrations. We provide the full mathe…

Imitation Learning

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic

2026-05-09 · Jiamin He, Samuel Neumann, Jincheng Mei, Adam White 외 arxiv

Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefits of this complexity remain elusive. Mixture policies are notably abse…

Reinforcement Learning