paper-with-me

홈 › Papers

A CMDP-within-online framework for Meta-Safe Reinforcement Learning

2024-05-26 · Vanshaj Khattar, Yuhao Ding, Bilgehan Sel, Javad Lavaei, Ming Jin

Meta-reinforcement learning has widely been used as a learning-to-learn framework to solve unseen tasks with limited experience. However, the aspect of constraint violations has not been adequately addressed in the existing works, making their application restricted in real-world settings. In this paper, we study the problem of meta-safe reinforcement learning (Meta-SRL) through the CMDP-within-online framework to establish the first provable guarantees in this important setting. We obtain task-averaged regret bounds for the reward maximization (optimality gap) and constraint violations using gradient-based meta-learning and show that the task-averaged optimality gap and constraint satisfaction improve with task-similarity in a static environment or task-relatedness in a dynamic environment. Several technical challenges arise when making this framework practical. To this end, we propose a meta-algorithm that performs inexact online learning on the upper bounds of within-task optimality gap and constraint violations estimated by off-policy stationary distribution corrections. Furthermore, we enable the learning rates to be adapted for every task and extend our approach to settings with a competing dynamically changing oracle. Finally, experiments are conducted to demonstrate the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2405.16601

Code (0)

등록된 구현이 없습니다.

Tasks

Meta-LearningMeta Reinforcement Learningreinforcement-learningReinforcement LearningSafe Reinforcement Learning

Similar Papers 제목 키워드 기반

Seeing Before Colliding: Anticipatory Safe RL with Frozen Vision-Language Models

2026-06-09 · Samuel Tetteh, Cody Fleming arxiv

The cost signal that constrained-RL algorithms optimize against is almost always reactive: the simulator emits a non-zero cost only after a collision has begun, and the Lagrange multiplier of PPO-Lagrangian grows only af…

Threshold UCT: Cost-Constrained Monte Carlo Tree Search with Pareto Curves

2024-12-18 · Martin Kurečka, Václav Nevyhoštěný, Petr Novotný, Vít Unčovský

Constrained Markov decision processes (CMDPs), in which the agent optimizes expected payoffs while keeping the expected cost below a given threshold, are the leading framework for safe sequential decision making under st…

Decision MakingSequential Decision Making

Truly No-Regret Learning in Constrained MDPs

2024-02-24 · Adrian Müller, Pragnya Alatur, Volkan Cevher, Giorgia Ramponi 외

Constrained Markov decision processes (CMDPs) are a common way to model safety constraints in reinforcement learning. State-of-the-art methods for efficiently solving CMDPs are based on primal-dual algorithms. For these …

Safe Exploration by Solving Early Terminated MDP

2021-07-09 · Hao Sun, Ziping Xu, Meng Fang, Zhenghao Peng 외

Safe exploration is crucial for the real-world application of reinforcement learning (RL). Previous works consider the safe exploration problem as Constrained Markov Decision Process (CMDP), where the policies are being …

Reinforcement Learning (RL)Safe Exploration

Safe Wasserstein Constrained Deep Q-Learning

2020-02-07 · Aaron Kandel, Scott J. Moura

This paper presents a distributionally robust Q-Learning algorithm (DrQ) which leverages Wasserstein ambiguity sets to provide idealistic probabilistic out-of-sample safety guarantees during online learning. First, we fo…

Q-Learning