paper-with-me

Papers

On the Generalization and Robustness in Conditional Value-at-Risk

2026-02-20 · Dinesh Karthik Mulumudi, Piyushi Manupriya, Gholamali Aminian, Anant Raj arxiv

Conditional Value-at-Risk (CVaR) is a widely used risk-sensitive objective for learning under rare but high-impact losses, yet its statistical behavior under heavy-tailed data remains poorly understood. Unlike expectation-based risk, CVaR depends on an endogenous, data-dependent quantile, which couples tail averaging with threshold estimation and fundamentally alters both generalization and robustness properties. In this work, we develop a learning-theoretic analysis of CVaR-based empirical risk minimization under heavy-tailed and contaminated data. We establish sharp, high-probability generalization and excess risk bounds under minimal moment assumptions, covering fixed hypotheses, finite and infinite classes, and extending to $β$-mixing dependent data; we further show that these rates are minimax optimal. To capture the intrinsic quantile sensitivity of CVaR, we derive a uniform Bahadur-Kiefer type expansion that isolates a threshold-driven error term absent in mean-risk ERM and essential in heavy-tailed regimes. We complement these results with robustness guarantees by proposing a truncated median-of-means CVaR estimator that achieves optimal rates under adversarial contamination. Finally, we show that CVaR decisions themselves can be intrinsically unstable under heavy tails, establishing a fundamental limitation on decision robustness even when the population optimum is well separated. Together, our results provide a principled characterization of when CVaR learning generalizes and is robust, and when instability is unavoidable due to tail scarcity.

📄 PDF Abstract BibTeX arXiv:2602.18053

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training

2026-02-05 · Dingwei Zhu, Zhiheng Xi, Shihan Dou, Jiahan Li 외 arxiv

Training reinforcement learning (RL) systems in real-world environments remains challenging due to noisy supervision and poor out-of-domain (OOD) generalization, especially in LLM post-training. Recent distributional RL …

Reinforcement Learning

STL Robustness Risk over Discrete-Time Stochastic Processes

2021-04-03 · Lars Lindemann, Nikolai Matni, George J. Pappas

We present a framework to interpret signal temporal logic (STL) formulas over discrete-time stochastic processes in terms of the induced risk. Each realization of a stochastic process either satisfies or violates an STL …

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training

2025-12-03 · Dingwei Zhu, Zhiheng Xi, Shihan Dou, Yuhui Wang 외 arxiv

Reinforcement learning (RL) has shown strong performance in LLM post-training, but real-world deployment often involves noisy or incomplete supervision. In such settings, complex and unreliable supervision signals can de…

Reinforcement Learning

Towards Safe Reinforcement Learning via Constraining Conditional Value at Risk

2021-06-18 · ICML Workshop AML 2021 7 · Chengyang Ying, Xinning Zhou, Dong Yan, Jun Zhu

Though deep reinforcement learning (DRL) has obtained substantial success, it may encounter catastrophic failures due to the intrinsic uncertainty caused by stochastic policies and environment variability. To address thi…

continuous-controlContinuous ControlDeep Reinforcement LearningMuJoCo+4

Risk-Averse Reinforcement Learning via Dynamic Time-Consistent Risk Measures

2023-01-14 · Xian Yu, Siqian Shen

Traditional reinforcement learning (RL) aims to maximize the expected total reward, while the risk of uncertain outcomes needs to be controlled to ensure reliable performance in a risk-averse setting. In this paper, we c…

Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)