Safe Exploration by Solving Early Terminated MDP
Safe exploration is crucial for the real-world application of reinforcement learning (RL). Previous works consider the safe exploration problem as Constrained Markov Decision Process (CMDP), where the policies are being optimized under constraints. However, when encountering any potential dangers, human tends to stop immediately and rarely learns to behave safely in danger. Motivated by human learning, we introduce a new approach to address safe RL problems under the framework of Early Terminated MDP (ET-MDP). We first define the ET-MDP as an unconstrained MDP with the same optimal value function as its corresponding CMDP. An off-policy algorithm based on context models is then proposed to solve the ET-MDP, which thereby solves the corresponding CMDP with better asymptotic performance and improved learning efficiency. Experiments on various CMDP tasks show a substantial improvement over previous methods that directly solve CMDP.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement Learning (RL)Safe ExplorationSimilar Papers 제목 키워드 기반
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Automated red-teaming has become a crucial approach for uncovering vulnerabilities in large language models (LLMs). However, most existing methods focus on isolated safety flaws, limiting their ability to adapt to dynami…
Red TeamingAnytime Safe Reinforcement Learning
This paper considers the problem of solving constrained reinforcement learning problems with anytime guarantees, meaning that the algorithmic solution returns a safe policy regardless of when it is terminated. Drawing in…
reinforcement-learningReinforcement LearningSafe Reinforcement LearningSafe Exploration Using Bayesian World Models and Log-Barrier Optimization
A major challenge in deploying reinforcement learning in online tasks is ensuring that safety is maintained throughout the learning process. In this work, we propose CERL, a new method for solving constrained Markov deci…
Safe ExplorationESO Valuation with Job Termination Risk and Jumps in Stock Price
Employee stock options (ESOs) are American-style call options that can be terminated early due to employment shock. This paper studies an ESO valuation framework that accounts for job termination risk and jumps in the co…
Improving RL Exploration for LLM Reasoning through Retrospective Replay
Reinforcement learning (RL) has increasingly become a pivotal technique in the post-training of large language models (LLMs). The effective exploration of the output space is essential for the success of RL. We observe t…
Code GenerationMathematical ReasoningReinforcement Learning (RL)