paper-with-me

Papers

Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control

2026-03-11 · Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee, Scott Niekum arxiv

Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events. This limitation is problematic when robustness and risk sensitivity are critical. Stochastic dominance offers a principled alternative by comparing entire cost distributions rather than just their averages, enabling direct control over tail risks and potential out-of-distribution failures that expectation-based constraints may overlook. In this work, we propose Risk-sensitive Alignment via Dominance (RAD), a novel alignment framework that replaces scalar expected cost constraints with First-Order Stochastic Dominance (FSD) constraints. We operationalize this constraint by comparing the target policy's cost distribution to that of a reference policy within an Optimal Transport (OT) framework, using entropic regularization and Sinkhorn iterations to obtain a differentiable and computationally efficient objective for stable end-to-end optimization. Furthermore, we introduce quantile-weighted FSD constraints and show that weighted FSD universally controls a broad class of Spectral Risk Measures (SRMs), so that improvements under weighted dominance imply guaranteed improvements in the corresponding spectral risk. This provides a principled mechanism for tuning a model's risk profile via the quantile weighting function. Empirical results demonstrate that RAD improves harmlessness over baselines while remaining competitive in helpfulness, and exhibits greater robustness on out-of-distribution harmlessness evaluations.

📄 PDF Abstract BibTeX arXiv:2603.10938

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Beyond Expectations: Learning with Stochastic Dominance Made Practical

2024-02-05 · Shicong Cen, Jincheng Mei, Hanjun Dai, Dale Schuurmans 외

Stochastic dominance models risk-averse preferences for decision making with uncertain outcomes, which naturally captures the intrinsic structure of the underlying uncertainty, in contrast to simply resorting to the expe…

Decision MakingPortfolio Optimization

Robust Statistical Comparison of Random Variables with Locally Varying Scale of Measurement

2023-06-22 · Christoph Jansen, Georg Schollmeyer, Hannah Blocher, Julian Rodemann 외

Spaces with locally varying scale of measurement, like multidimensional structures with differently scaled dimensions, are pretty common in statistics and machine learning. Nevertheless, it is still understood as an open…

Almost Dominance: Inference and Application

2023-12-04 · Xiaojun Song, Zhenting Sun

This paper proposes a general framework for inference on three types of almost dominances: Almost Lorenz dominance, almost inverse stochastic dominance, and almost stochastic dominance. We first generalize almost Lorenz …

Safe RLHF: Safe Reinforcement Learning from Human Feedback

2023-10-19 · Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 외

With the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tension between the objectives of helpfulness …

reinforcement-learningReinforcement LearningSafe Reinforcement Learning

Safety Third: Roy's Criterion and Higher Order Moments

2015-06-13

Roy's `Safety First' criterion for selecting one risky asset from many is adapted to the case of non-normal returns, via Cornish Fisher expansion. The resulting investment objective is consistent with first order stochas…