Effects of Safety State Augmentation on Safe Exploration
Safe exploration is a challenging and important problem in model-free reinforcement learning (RL). Often the safety cost is sparse and unknown, which unavoidably leads to constraint violations -- a phenomenon ideally to be avoided in safety-critical applications. We tackle this problem by augmenting the state-space with a safety state, which is nonnegative if and only if the constraint is satisfied. The value of this state also serves as a distance toward constraint violation, while its initial value indicates the available safety budget. This idea allows us to derive policies for scheduling the safety budget during training. We call our approach Simmer (Safe policy IMproveMEnt for RL) to reflect the careful nature of these schedules. We apply this idea to two safe RL problems: RL with constraints imposed on an average cost, and RL with constraints imposed on a cost with probability one. Our experiments suggest that "simmering, a safe algorithm can improve safety during training for both settings. We further show that Simmer can stabilize training and improve the performance of safe RL with average constraints.
Code (1)
Tasks
Reinforcement Learning (RL)Safe ExplorationSchedulingSimilar Papers 제목 키워드 기반
Safety Representations for Safer Policy Learning
Reinforcement learning algorithms typically necessitate extensive exploration of the state space to find optimal policies. However, in safety-critical applications, the risks associated with such exploration can lead to …
Safe ExplorationProbabilistic Counterexample Guidance for Safer Reinforcement Learning (Extended Version)
Safe exploration aims at addressing the limitations of Reinforcement Learning (RL) in safety-critical scenarios, where failures during trial-and-error learning may incur high costs. Several methods exist to incorporate e…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe ExplorationSystem III: Learning with Domain Knowledge for Safety Constraints
Reinforcement learning agents naturally learn from extensive exploration. Exploration is costly and can be unsafe in $\textit{safety-critical}$ domains. This paper proposes a novel framework for incorporating domain know…
Safe ExplorationParenting: Safe Reinforcement Learning from Human Input
Autonomous agents trained via reinforcement learning present numerous safety concerns: reward hacking, negative side effects, and unsafe exploration, among others. In the context of near-future autonomous agents, operati…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement LearningSafety-Guided Deep Reinforcement Learning via Online Gaussian Process Estimation
An important facet of reinforcement learning (RL) has to do with how the agent goes about exploring the environment. Traditional exploration strategies typically focus on efficiency and ignore safety. However, for practi…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1