Model-Based Uncertainty in Value Functions
We consider the problem of quantifying uncertainty over expected cumulative rewards in model-based reinforcement learning. In particular, we focus on characterizing the variance over values induced by a distribution over MDPs. Previous work upper bounds the posterior variance over values by solving a so-called uncertainty Bellman equation, but the over-approximation may result in inefficient exploration. We propose a new uncertainty Bellman equation whose solution converges to the true posterior variance over values and explicitly characterizes the gap in previous work. Moreover, our uncertainty quantification technique is easily integrated into common exploration strategies and scales naturally beyond the tabular setting by using standard deep reinforcement learning architectures. Experiments in difficult exploration tasks, both in tabular and continuous control settings, show that our sharper uncertainty estimates improve sample-efficiency.
Code (1)
Tasks
continuous-controlContinuous ControlDeep Reinforcement LearningmodelModel-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Uncertainty QuantificationSimilar Papers 제목 키워드 기반
Towards interval uncertainty propagation control in bivariate aggregation processes and the introduction of width-limited interval-valued overlap functions
Overlap functions are a class of aggregation functions that measure the overlapping degree between two values. Interval-valued overlap functions were defined as an extension to express the overlapping of interval-valued …
RelationDiverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning
Offline Reinforcement Learning (RL) faces distributional shift and unreliable value estimation, especially for out-of-distribution (OOD) actions. To address this, existing uncertainty-based methods penalize the value fun…
DiversityReinforcement Learning (RL)Uncertainty QuantificationEpistemic Deep Learning
The belief function approach to uncertainty quantification as proposed in the Demspter-Shafer theory of evidence is established upon the general mathematical models for set-valued observations, called random sets. Set-va…
Deep LearningUncertainty QuantificationQuantifying and Computing Covariance Uncertainty
In this work, we consider the problem of bounding the values of a covariance function corresponding to a continuous-time stationary stochastic process or signal. Specifically, for two signals whose covariance functions a…
Short time quaternion quadratic phase Fourier transform and its uncertainty principles
In this paper, we extend the quadratic phase Fourier transform of a complex valued functions to that of the quaternion valued functions of two variables. We call it the quaternion quadratic phase Fourier transform (QQPFT…
Relation