paper-with-me

홈 › Papers

Negative impact of heavy-tailed uncertainty and error distributions on the reliability of calibration statistics for machine learning regression tasks

2024-02-15 · Pascal Pernot

Average calibration of the (variance-based) prediction uncertainties of machine learning regression tasks can be tested in two ways: one is to estimate the calibration error (CE) as the difference between the mean absolute error (MSE) and the mean variance (MV); the alternative is to compare the mean squared z-scores (ZMS) to 1. The problem is that both approaches might lead to different conclusions, as illustrated in this study for an ensemble of datasets from the recent machine learning uncertainty quantification (ML-UQ) literature. It is shown that the estimation of MV, MSE and their confidence intervals becomes unreliable for heavy-tailed uncertainty and error distributions, which seems to be a frequent feature of ML-UQ datasets. By contrast, the ZMS statistic is less sensitive and offers the most reliable approach in this context, still acknowledging that datasets with heavy-tailed z-scores distributions should be considered with great care. Unfortunately, the same problem is expected to affect also conditional calibrations statistics, such as the popular ENCE, and very likely post-hoc calibration methods based on similar statistics. Several solutions to circumvent the outlined problems are proposed.

📄 PDF Abstract BibTeX arXiv:2402.10043

Code (1)

ppernot/2024_rce 공식 구현

Tasks

regressionUncertainty Quantification

Similar Papers 제목 키워드 기반

Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification

2024-07-19 · Thomas Kwa, Drake Thomas, Adrià Garriga-Alonso

When applying reinforcement learning from human feedback (RLHF), the reward is learned from data and, therefore, always has some error. It is common to mitigate this by regularizing the policy with KL divergence from a b…

Maximum Likelihood Uncertainty Estimation: Robustness to Outliers

2022-02-03 · Deebul S. Nair, Nico Hochgeschwender, Miguel A. Olivares-Mendez

We benchmark the robustness of maximum likelihood based uncertainty estimation methods to outliers in training data for regression tasks. Outliers or noisy labels in training data results in degraded performances as well…

Depth EstimationMonocular Depth Estimationregression

Best-of-Both-Worlds for Heavy-Tailed Markov Decision Processes

2026-02-01 · Yu Chen, Yuhao Liu, Jiatai Huang, Yihan Du 외 arxiv

We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing approaches for HTMDPs are conservative in stochastic environments and lack adaptivity in adversarial regimes. In this work, we…

Robust Offline Reinforcement learning with Heavy-Tailed Rewards

2023-10-28 · Jin Zhu, Runzhe Wan, Zhengling Qi, Shikai Luo 외

This paper endeavors to augment the robustness of offline reinforcement learning (RL) in scenarios laden with heavy-tailed rewards, a prevalent circumstance in real-world applications. We propose two algorithmic framewor…

Offline RLOff-policy evaluationreinforcement-learningReinforcement Learning+1

Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise

2026-01-27 · Hongxu Chen, Ke Wei, Xiaoming Yuan, Luo Luo arxiv

The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance…

Stochastic Optimization