paper-with-me

홈 › Papers

General non-linear Bellman equations

2019-07-08 · Hado van Hasselt, John Quan, Matteo Hessel, Zhongwen Xu, Diana Borsa, Andre Barreto

We consider a general class of non-linear Bellman equations. These open up a design space of algorithms that have interesting properties, which has two potential advantages. First, we can perhaps better model natural phenomena. For instance, hyperbolic discounting has been proposed as a mathematical model that matches human and animal data well, and can therefore be used to explain preference orderings. We present a different mathematical model that matches the same data, but that makes very different predictions under other circumstances. Second, the larger design space can perhaps lead to algorithms that perform better, similar to how discount factors are often used in practice even when the true objective is undiscounted. We show that many of the resulting Bellman operators still converge to a fixed point, and therefore that the resulting algorithms are reasonable and inherit many beneficial properties of their linear counterparts.

📄 PDF Abstract BibTeX arXiv:1907.03687

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Deep Learning for High-Dimensional Hamilton-Jacobi-Bellman Equations

2019-07-11 · Tenavi Nakamura-Zimmerer, Qi Gong, Wei Kang

Computing optimal feedback controls for nonlinear systems generally requires solving Hamilton-Jacobi-Bellman (HJB) equations, which are notoriously difficult when the state dimension is large. Existing strategies for hig…

Deep LearningvalidVocal Bursts Intensity Prediction

On solutions of the distributional Bellman equation

2022-01-31 · Julian Gerstenberg, Ralph Neininger, Denis Spiegel

In distributional reinforcement learning not only expected returns but the complete return distributions of a policy are taken into account. The return distribution for a fixed policy is given as the solution of an assoc…

Distributional Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Deep Graphic FBSDEs for Opinion Dynamics Stochastic Control

2022-04-05 · Tianrong Chen, Ziyi Wang, Evangelos A. Theodorou

In this paper, we present a scalable deep learning approach to solve opinion dynamics stochastic optimal control problems with mean field term coupling in the dynamics and cost function. Our approach relies on the probab…

LEMMA

Forward and Backward Bellman equations improve the efficiency of EM algorithm for DEC-POMDP

2021-03-19 · Takehiro Tottori, Tetsuya J. Kobayashi

Decentralized partially observable Markov decision process (DEC-POMDP) models sequential decision making problems by a team of agents. Since the planning of DEC-POMDP can be interpreted as the maximum likelihood estimati…

Computational EfficiencyDecision MakingSequential Decision Making

Bellman-consistent Pessimism for Offline Reinforcement Learning

2021-06-13 · NeurIPS 2021 12 · Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro 외

The use of pessimism, when reasoning about datasets lacking exhaustive exploration has recently gained prominence in offline reinforcement learning. Despite the robustness it adds to the algorithm, overly pessimistic rea…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)