paper-with-me

홈 › Papers

Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk

2026-01-30 · Rohan Tangri, Jan-Peter Calliess arxiv

We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems. We employ Cantelli's inequality to obtain a tractable, conservative and smooth bound on the VaR constraint based on the first two moments of the cost return. This yields a constraint estimator that remains stable with tight violation thresholds in dense cost regimes. Extending the trust-region framework of the Constrained Policy Optimization (CPO) method, we further provide worst-case bounds for both policy improvement and constraint violation during the training process. Empirically, across continuous-control safety benchmarks, Canary most reliably satisfies its constraint, with the fewest violations and the earliest permanent satisfaction, while remaining reward-competitive with other baselines that also satisfy.

📄 PDF Abstract BibTeX arXiv:2601.22993

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Risk-Sensitive Deep RL: Variance-Constrained Actor-Critic Provably Finds Globally Optimal Policy

2020-12-28 · Han Zhong, Xun Deng, Ethan X. Fang, Zhuoran Yang 외

While deep reinforcement learning has achieved tremendous successes in various applications, most existing works only focus on maximizing the expected value of total return and thus ignore its inherent stochasticity. Suc…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Chebyshev-Cantelli PAC-Bayes-Bennett Inequality for the Weighted Majority Vote

2021-06-25 · NeurIPS 2021 12 · Yi-Shan Wu, Andrés R. Masegosa, Stephan S. Lorenzen, Christian Igel 외

We present a new second-order oracle bound for the expected risk of a weighted majority vote. The bound is based on a novel parametric form of the Chebyshev- Cantelli inequality (a.k.a. one-sided Chebyshev's), which is a…

Form

Bounded Ratio Reinforcement Learning

2026-04-20 · Yunke Ao, Le Chen, Bruce D. Lee, Assefa S. Wahd 외 arxiv

Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However, there is a significant disconnect betw…

Reinforcement Learning

The Exact Worst-Case Tail Probability under Bounded Kurtosis

2026-07-06 · Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao arxiv

We determine exactly what a kurtosis bound buys for one-sided tail control. For the class $\mathcal{C}(κ)$ of real random variables with mean $0$, variance $1$, and fourth moment at most $κ$, the skewness left free, we c…

Submultiplicative Glivenko-Cantelli and Uniform Convergence of Revenues

2017-05-23 · NeurIPS 2017 12 · Noga Alon, Moshe Babaioff, Yannai A. Gonczarowski, Yishay Mansour 외

In this work we derive a variant of the classic Glivenko-Cantelli Theorem, which asserts uniform convergence of the empirical Cumulative Distribution Function (CDF) to the CDF of the underlying distribution. Our variant …