paper-with-me

Papers

AutoCost: Evolving Intrinsic Cost for Zero-violation Reinforcement Learning

2023-01-24 · Tairan He, WeiYe Zhao, Changliu Liu

Safety is a critical hurdle that limits the application of deep reinforcement learning (RL) to real-world control tasks. To this end, constrained reinforcement learning leverages cost functions to improve safety in constrained Markov decision processes. However, such constrained RL methods fail to achieve zero violation even when the cost limit is zero. This paper analyzes the reason for such failure, which suggests that a proper cost function plays an important role in constrained RL. Inspired by the analysis, we propose AutoCost, a simple yet effective framework that automatically searches for cost functions that help constrained RL to achieve zero-violation performance. We validate the proposed method and the searched cost function on the safe RL benchmark Safety Gym. We compare the performance of augmented agents that use our cost function to provide additive intrinsic costs with baseline agents that use the same policy learners but with only extrinsic costs. Results show that the converged policies with intrinsic costs in all environments achieve zero constraint violation and comparable performance with baselines.

📄 PDF Abstract BibTeX arXiv:2301.10339

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Controlling Underestimation Bias in Constrained Reinforcement Learning for Safe Exploration

2026-01-17 · Shiqing Gao, Jiaxin Ding, Luoyi Fu, Xinbing Wang arxiv

Constrained Reinforcement Learning (CRL) aims to maximize cumulative rewards while satisfying constraints. However, existing CRL algorithms often encounter significant constraint violations during training, limiting thei…

Reinforcement Learning

Extreme Value Policy Optimization for Safe Reinforcement Learning

2026-01-17 · Shiqing Gao, Yihang Zhou, Shuai Shao, Haoyu Luo 외 arxiv

Ensuring safety is a critical challenge in applying Reinforcement Learning (RL) to real-world scenarios. Constrained Reinforcement Learning (CRL) addresses this by maximizing returns under predefined constraints, typical…

Reinforcement Learning

Safe Online Convex Optimization with Multi-Point Feedback

2024-07-16 · Spencer Hutchinson, Mahnoosh Alizadeh

Motivated by the stringent safety requirements that are often present in real-world applications, we study a safe online convex optimization setting where the player needs to simultaneously achieve sublinear regret and z…

Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time Violations

2021-08-04 · NeurIPS 2021 12 · Yuping Luo, Tengyu Ma

Training-time safety violations have been a major concern when we deploy reinforcement learning algorithms in the real world. This paper explores the possibility of safe RL algorithms with zero training-time safety viola…

reinforcement-learningReinforcement Learning (RL)Safe Reinforcement Learning

Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings

2024-12-18 · Enming Luo, Wei Qiao, Katie Warren, Jingxiang Li 외

We present a scalable and agile approach for ads image content moderation at Google, addressing the challenges of moderating massive volumes of ads with diverse content and evolving policies. The proposed method utilizes…

zero-shot-classificationZero-Shot Learning