paper-with-me

홈 › Papers

A general Markov decision process formalism for action-state entropy-regularized reward maximization

2023-02-02 · Dmytro Grytskyy, Jorge Ramírez-Ruiz, Rubén Moreno-Bote

Previous work has separately addressed different forms of action, state and action-state entropy regularization, pure exploration and space occupation. These problems have become extremely relevant for regularization, generalization, speeding up learning and providing robust solutions at unprecedented levels. However, solutions of those problems are hectic, ranging from convex and non-convex optimization, and unconstrained optimization to constrained optimization. Here we provide a general dual function formalism that transforms the constrained optimization problem into an unconstrained convex one for any mixture of action and state entropies. The cases with pure action entropy and pure state entropy are understood as limits of the mixture.

📄 PDF Abstract BibTeX arXiv:2302.01098

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Partially Observable History Process

2021-11-15 · Dustin Morrill, Amy R. Greenwald, Michael Bowling

We introduce the partially observable history process (POHP) formalism for reinforcement learning. POHP centers around the actions and observations of a single agent and abstracts away the presence of other players witho…

Formreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

2026-08-24 · Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi, Vijay Gupta 외 arxiv

Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism…

A Translation of Probabilistic Event Calculus into Markov Decision Processes

2025-07-17 · Lyris Xu, Fabio Aurelio D'Asaro, Luke Dickens

Probabilistic Event Calculus (PEC) is a logical framework for reasoning about actions and their effects in uncertain environments, which enables the representation of probabilistic narratives and computation of temporal …

Translation

Universal Decision Models

2021-10-28 · Sridhar Mahadevan

Humans are universal decision makers: we reason causally to understand the world; we act competitively to gain advantage in commerce, games, and war; and we are able to learn to make better decisions through trial and er…

Causal Inference

Modularity in Reinforcement Learning via Algorithmic Independence in Credit Assignment

2021-06-28 · ICLR Workshop Learning_to_Learn 2021 5 · Michael Chang, Sidhant Kaushik, Sergey Levine, Thomas L. Griffiths

Many transfer problems require re-using previously optimal decisions for solving new tasks, which suggests the need for learning algorithms that can modify the mechanisms for choosing certain actions independently of tho…

Decision MakingPolicy Gradient Methodsreinforcement-learningReinforcement Learning+1