paper-with-me

홈 › Papers

Projection-Based Constrained Policy Optimization

2020-10-07 · ICLR 2020 1 · Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. Ramadge

We consider the problem of learning control policies that optimize a reward function while satisfying constraints due to considerations of safety, fairness, or other costs. We propose a new algorithm, Projection-Based Constrained Policy Optimization (PCPO). This is an iterative method for optimizing policies in a two-step process: the first step performs a local reward improvement update, while the second step reconciles any constraint violation by projecting the policy back onto the constraint set. We theoretically analyze PCPO and provide a lower bound on reward improvement, and an upper bound on constraint violation, for each policy update. We further characterize the convergence of PCPO based on two different metrics: $\normltwo$ norm and Kullback-Leibler divergence. Our empirical results over several control tasks demonstrate that PCPO achieves superior performance, averaging more than 3.5 times less constraint violation and around 15\% higher reward compared to state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2010.03152

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Similar Papers 제목 키워드 기반

Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space

2026-02-24 · Wang Zixian arxiv

We present Group Orthogonalized Policy Optimization (GOPO), a new alignment algorithm for large language models derived from the geometry of Hilbert function spaces. Instead of optimizing on the probability simplex and i…

Mathematical Reasoning

Constrained Policy Optimization via Sampling-Based Weight-Space Projection

2025-12-15 · Shengfan Cao, Francesco Borrelli, Eunhyek Joa arxiv

Safety-critical learning requires policies that improve performance without leaving the safe operating regime. We study constrained policy learning where model parameters must satisfy rollout-based safety constraints tha…

Escaping from Zero Gradient: Revisiting Action-Constrained Reinforcement Learning via Frank-Wolfe Policy Optimization

2021-02-22 · Jyun-Li Lin, Wei Hung, Shang-Hsuan Yang, Ping-Chun Hsieh 외

Action-constrained reinforcement learning (RL) is a widely-used approach in various real-world applications, such as scheduling in networked systems with resource constraints and control of a robot with kinematic constra…

Reinforcement Learning (RL)Scheduling

Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI

2026-04-19 · Vinil Pasupuleti, Shyalendar Reddy Allala, Siva Rama Krishna Varma Bayyavarapu, Shrey Tyagi arxiv

Enterprise AI systems increasingly deploy multiple intelligent agents across mission-critical workflows that must satisfy hard policy constraints, bounded risk exposure, and comprehensive auditability (SOX, HIPAA, GDPR).…

Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space

2026-01-18 · Wang Zixian arxiv

We propose Orthogonalized Policy Optimization (OPO), a principled framework for large language model alignment derived from optimization in the Hilbert function space L2(pi_k). Lifting policy updates from the probability…