Multi-Objective Value Iteration with Parameterized Threshold-Based Safety Constraints
We consider an environment with multiple reward functions. One of them represents goal achievement and the others represent instantaneous safety conditions. We consider a scenario where the safety rewards should always be above some thresholds. The thresholds are parameters with values that differ between users. %The thresholds are not known at the time the policy is being designed. We efficiently compute a family of policies that cover all threshold-based constraints and maximize the goal achievement reward. We introduce a new parameterized threshold-based scalarization method of the reward vector that encodes our objective. We present novel data structures to store the value functions of the Bellman equation that allow their efficient computation using the value iteration algorithm. We present results for both discrete and continuous state spaces.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A multilevel thresholding algorithm using Electromagnetism Optimization
Segmentation is one of the most important tasks in image processing. It consist in classify the pixels into two or more groups depending on their intensity levels and a threshold value. The quality of the segmentation de…
Image SegmentationSegmentationSemantic SegmentationA Multi-objective Newton Optimization Algorithm for Hyper-Parameter Search
This study proposes a Newton based multiple objective optimization algorithm for hyperparameter search. The first order differential (gradient) is calculated using finite difference method and a gradient matrix with vect…
Bayesian Optimizationobject-detectionObject DetectionEnhanced Low-Rank Matrix Approximation
This letter proposes to estimate low-rank matrices by formulating a convex optimization problem with non-convex regularization. We employ parameterized non-convex penalty functions to estimate the non-zero singular value…
DenoisingImage DenoisingThe association problem in wireless networks: a Policy Gradient Reinforcement Learning approach
The purpose of this paper is to develop a self-optimized association algorithm based on PGRL (Policy Gradient Reinforcement Learning), which is both scalable, stable and robust. The term robust means that performance deg…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Mirror Natural Evolution Strategies
The zeroth-order optimization has been widely used in machine learning applications. However, the theoretical study of the zeroth-order optimization focus on the algorithms which approximate (first-order) gradients using…