paper-with-me

홈 › Papers

Conformal Tail Risk Control for Large Language Model Alignment

2025-02-27 · Catherine Yu-Chi Chen, Jingyan Shen, Zhun Deng, Lihua Lei

Recent developments in large language models (LLMs) have led to their widespread usage for various tasks. The prevalence of LLMs in society implores the assurance on the reliability of their performance. In particular, risk-sensitive applications demand meticulous attention to unexpectedly poor outcomes, i.e., tail events, for instance, toxic answers, humiliating language, and offensive outputs. Due to the costly nature of acquiring human annotations, general-purpose scoring models have been created to automate the process of quantifying these tail events. This phenomenon introduces potential human-machine misalignment between the respective scoring mechanisms. In this work, we present a lightweight calibration framework for blackbox models that ensures the alignment of humans and machines with provable guarantees. Our framework provides a rigorous approach to controlling any distortion risk measure that is characterized by a weighted average of quantiles of the loss incurred by the LLM with high confidence. The theoretical foundation of our method relies on the connection between conformal risk control and a traditional family of statistics, i.e., L-statistics. To demonstrate the utility of our framework, we conduct comprehensive experiments that address the issue of human-machine misalignment.

📄 PDF Abstract BibTeX arXiv:2502.20285

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal Inference

2026-05-29 · Catherine Chen, Jingyan Shen, Zhun Deng, Lihua Lei arxiv

We present an online, distribution-free framework for controlling the Conditional Value-at-Risk (CVaR), extending conformal tail risk control to non-stationary and adversarial environments. Unlike classical risk control …

Conformal Risk Training: End-to-End Optimization of Conformal Risk Control

2025-10-09 · Christopher Yeh, Nicolas Christianson, Adam Wierman, Yisong Yue arxiv

While deep learning models often achieve high predictive accuracy, their predictions typically do not come with any provable guarantees on risk or reliability, which are critical for deployment in high-stakes application…

Conformal Risk Control

2022-08-04 · Anastasios N. Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei 외

We extend conformal prediction to control the expected value of any monotone loss function. The algorithm generalizes split conformal prediction together with its coverage guarantee. Like conformal prediction, the confor…

Conformal PredictionPrediction

Selective Conformal Risk Control

2025-12-14 · Yunpeng Xu, Wenge Guo, Zhi Wei arxiv

Reliable uncertainty quantification is essential for deploying machine learning systems in high-stakes domains. Conformal prediction provides distribution-free coverage guarantees but often produces overly large predicti…

Conformal Risk Control for Ordinal Classification

2024-05-01 · Yunpeng Xu, Wenge Guo, Zhi Wei

As a natural extension to the standard conformal prediction method, several conformal risk control methods have been recently developed and applied to various learning problems. In this work, we seek to control the confo…

ClassificationConformal PredictionDiabetic Retinopathy DetectionOrdinal Classification