An Active Parameter Learning Approach to The Identification of Safe Regions
We consider the problem of identification of safe regions in the environment of an autonomous system. The environment is divided into a finite collections of Voronoi cells, with each cell having a representative, the Voronoi center. The extent to which each region is considered to be safe by an oracle is captured through a trust distribution. The trust placed by the oracle conditioned on the region is modeled through a Bernoulli distribution whose the parameter depends on the region. The parameters are unknown to the system. However, if the agent were to visit a given region, it will receive a binary valued random response from the oracle on whether the oracle trusts the region or not. The objective is to design a path for the agent where, by traversing through the centers of the cells, the agent is eventually able to label each cell safe or unsafe. To this end, we formulate an active parameter learning problem with the objective of minimizing visits or stays in potentially unsafe regions. The active learning problem is formulated as a finite horizon stochastic control problem where the cost function is derived utilizing the large deviations principle (LDP). The challenges associated with a dynamic programming approach to solve the problem are analyzed. Subsequently, the optimization problem is relaxed to obtain single-step optimization problems for which closed form solution is obtained. Using the solution, we propose an algorithm for the active learning of the parameters. A relationship between the trust distributions and the label of a cell is defined and subsequently a classification algorithm is proposed to identify the safe regions. We prove that the algorithm identifies the safe regions with finite number of visits to unsafe regions. We demonstrate the algorithm through an example.
Code (0)
등록된 구현이 없습니다.
Tasks
Active LearningSimilar Papers 제목 키워드 기반
Can LLM Safety Be Ensured by Constraining Parameter Regions?
Large language models (LLMs) are often assumed to contain ``safety regions'' -- parameter subsets whose modification directly influences safety behaviors. We conduct a systematic evaluation of four safety region identifi…
Learning safety critics via a non-contractive binary bellman operator
The inability to naturally enforce safety in Reinforcement Learning (RL), with limited failures, is a core challenge impeding its use in real-world applications. One notion of safety of vast practical relevance is the ab…
Reinforcement Learning (RL)Constraining the Parameters of High-Dimensional Models with Active Learning
Constraining the parameters of physical models with $>5-10$ parameters is a widespread problem in fields like particle physics and astronomy. The generation of data to explore this parameter space often requires large am…
Active LearningAstronomyVocal Bursts Intensity PredictionActive exploration in adaptive model predictive control
A dual adaptive model predictive control (MPC) algorithm is presented for linear, time-invariant systems subject to bounded disturbances and parametric uncertainty in the state-space matrices. Online set-membership ident…
modelModel Predictive ControlUnscented Bayesian Optimization for Safe Robot Grasping
We address the robot grasp optimization problem of unknown objects considering uncertainty in the input space. Grasping unknown objects can be achieved by using a trial and error exploration strategy. Bayesian optimizati…
Bayesian Optimization