paper-with-me

홈 › Papers

Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning

2023-10-05 · NeurIPS 2023 11 · Yihang Yao, Zuxin Liu, Zhepeng Cen, Jiacheng Zhu, Wenhao Yu, Tingnan Zhang, Ding Zhao

Safe reinforcement learning (RL) focuses on training reward-maximizing agents subject to pre-defined safety constraints. Yet, learning versatile safe policies that can adapt to varying safety constraint requirements during deployment without retraining remains a largely unexplored and challenging area. In this work, we formulate the versatile safe RL problem and consider two primary requirements: training efficiency and zero-shot adaptation capability. To address them, we introduce the Conditioned Constrained Policy Optimization (CCPO) framework, consisting of two key modules: (1) Versatile Value Estimation (VVE) for approximating value functions under unseen threshold conditions, and (2) Conditioned Variational Inference (CVI) for encoding arbitrary constraint thresholds during policy optimization. Our extensive experiments demonstrate that CCPO outperforms the baselines in terms of safety and task performance while preserving zero-shot adaptation capabilities to different constraint thresholds data-efficiently. This makes our approach suitable for real-world dynamic applications.

📄 PDF Abstract BibTeX arXiv:2310.03718

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement LearningVariational Inference

Methods 이 논문이 사용한 방법론

Variational Inference 설명 없음

Similar Papers 제목 키워드 기반

Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning

2026-03-08 · Janaka Chathuranga Brahmanage, Akshat Kumar arxiv

Sequential decision making using Markov Decision Process underpins many realworld applications. Both model-based and model free methods have achieved strong results in these settings. However, real-world tasks must balan…

Reinforcement LearningDecision Making

Learning to explore when mistakes are not allowed

2025-02-19 · Charly Pecqueux-Guézénec, Stéphane Doncieux, Nicolas Perrin-Gilbert

Goal-Conditioned Reinforcement Learning (GCRL) provides a versatile framework for developing unified controllers capable of handling wide ranges of tasks, exploring environments, and adapting behaviors. However, its reli…

Safe ExplorationSafe Reinforcement Learning

Towards General Language-Conditioned Latent Safety Filters

2026-07-31 · Ihab Tabbara, Yuxuan Yang, Hussein Sibai arxiv

Robot policies are becoming increasingly general, with vision-language-action (VLA) models enabling a single policy to execute diverse tasks specified in natural language. Safe deployment, however, requires adapting not …

T-GMP: Terrain-conditioned Generative Motion Priors for Versatile and Natural Humanoid Locomotion

2026-06-05 · Junhong Guo, Hao Hu, Chen Chen, Haoxuan Han 외 arxiv

Achieving both anthropomorphic naturalness and robust terrain traversal remains a fundamental challenge in humanoid locomotion. Existing Reinforcement Learning (RL) approaches typically rely on fixed motion priors, limit…

Reinforcement Learning

Constrained Group Relative Policy Optimization

2026-02-05 · Roger Girgis, Rodrigue de Schaetzen, Luke Rowe, Azalée Robitaille 외 arxiv

Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with constrained policy optimization (e.g. for safety-critical domains) has not bee…