paper-with-me

Papers

Utility-Constrained Policy Optimization

2026-06-12 · Mehrdad Moghimi, Bernardo Avila Pires arxiv

Constrained MDPs (CMDPs) are a widely adopted framework for incorporating safety into RL agents; however, the framework does not support risk-sensitive constraints. This can be problematic: For example, CMDPs allow for optimal solutions that, in order to satisfy the risk-neutral constraints, mix infrequent catastrophic behaviors and frequent, overly conservative ones. Moreover, prior empirical results suggest that enforcing stricter, risk-sensitive constraints can improve performance even under risk-neutral evaluation. The natural framework to incorporate risk-sensitive constraints is utility-constrained MDPs (UCMDPs), but no practical solutions for this problem existed. In this work, we introduce a simple yet powerful methodology for UCMDPs and constrained RL. Besides allowing for risk-sensitive constraints, our framework does not require us to fix constraint limits in advance of training the agent, provided that a sensible range is known. This increases policy flexibility and, in practice, allows for adjustments to these limits at no extra training cost. Besides benefiting from the generality of the framework, our agent shows strong performance in practice, consistently matching or outperforming existing baselines in several Safety Gymnasium benchmark tasks.

📄 PDF Abstract BibTeX arXiv:2606.14029

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Entropic Risk Constrained Soft-Robust Policy Optimization

2020-06-20 · Reazul Hasan Russel, Bahram Behzadian, Marek Petrik

Having a perfect model to compute the optimal policy is often infeasible in reinforcement learning. It is important in high-stakes domains to quantify and manage risk induced by model uncertainties. Entropic risk measure…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI

2026-04-19 · Vinil Pasupuleti, Shyalendar Reddy Allala, Siva Rama Krishna Varma Bayyavarapu, Shrey Tyagi arxiv

Enterprise AI systems increasingly deploy multiple intelligent agents across mission-critical workflows that must satisfy hard policy constraints, bounded risk exposure, and comprehensive auditability (SOX, HIPAA, GDPR).…

A Deep Learning Based Resource Allocator for Communication Systems with Dynamic User Utility Demands

2023-11-08 · Pourya Behmandpoor, Mark Eisen, Panagiotis Patrinos, Marc Moonen

Deep learning (DL) based resource allocation (RA) has recently gained significant attention due to its performance efficiency. However, most related studies assume an ideal case where the number of users and their utilit…

The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

2026-06-02 · Xu Wan, Speed Zhu, Jianwei Cai, Guang Chen 외 arxiv

Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets. In this work, we formulate inference bud…

Continuous-time optimal investment with portfolio constraints: a reinforcement learning approach

2024-12-14 · Huy Chau, Duy Nguyen, Thai Nguyen

In a reinforcement learning (RL) framework, we study the exploratory version of the continuous time expected utility (EU) maximization problem with a portfolio constraint that includes widely-used financial regulations s…

Reinforcement Learning (RL)