General non-linear Bellman equations
We consider a general class of non-linear Bellman equations. These open up a design space of algorithms that have interesting properties, which has two potential advantages. First, we can perhaps better model natural phenomena. For instance, hyperbolic discounting has been proposed as a mathematical model that matches human and animal data well, and can therefore be used to explain preference orderings. We present a different mathematical model that matches the same data, but that makes very different predictions under other circumstances. Second, the larger design space can perhaps lead to algorithms that perform better, similar to how discount factors are often used in practice even when the true objective is undiscounted. We show that many of the resulting Bellman operators still converge to a fixed point, and therefore that the resulting algorithms are reasonable and inherit many beneficial properties of their linear counterparts.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Adaptive Deep Learning for High-Dimensional Hamilton-Jacobi-Bellman Equations
Computing optimal feedback controls for nonlinear systems generally requires solving Hamilton-Jacobi-Bellman (HJB) equations, which are notoriously difficult when the state dimension is large. Existing strategies for hig…
Deep LearningvalidVocal Bursts Intensity PredictionOn solutions of the distributional Bellman equation
In distributional reinforcement learning not only expected returns but the complete return distributions of a policy are taken into account. The return distribution for a fixed policy is given as the solution of an assoc…
Distributional Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Deep Graphic FBSDEs for Opinion Dynamics Stochastic Control
In this paper, we present a scalable deep learning approach to solve opinion dynamics stochastic optimal control problems with mean field term coupling in the dynamics and cost function. Our approach relies on the probab…
LEMMAForward and Backward Bellman equations improve the efficiency of EM algorithm for DEC-POMDP
Decentralized partially observable Markov decision process (DEC-POMDP) models sequential decision making problems by a team of agents. Since the planning of DEC-POMDP can be interpreted as the maximum likelihood estimati…
Computational EfficiencyDecision MakingSequential Decision MakingBellman-consistent Pessimism for Offline Reinforcement Learning
The use of pessimism, when reasoning about datasets lacking exhaustive exploration has recently gained prominence in offline reinforcement learning. Despite the robustness it adds to the algorithm, overly pessimistic rea…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)