paper-with-me

홈 › Papers

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

2026-02-13 · Sher Badshah, Ali Emami, Hassan Sajjad arxiv

Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose \textsc{Scope} (Selective Conformal Optimized Pairwise Evaluation), a framework that calibrates an acceptance threshold so that, under exchangeability, the error rate among non-abstained judgments is at most a user-specified level $α$. To supply \textsc{Scope} with a bias-neutral uncertainty signal, we introduce Bidirectional Preference Entropy (BPE), which queries the judge under both response positions and converts the order-averaged preference probability into an entropy-based score. Across various pairwise judging benchmarks, BPE outperforms standard confidence proxies in calibration and discrimination, while \textsc{Scope} consistently satisfies the target risk bound (empirical FDR $\approx 0.097$--$0.099$ at $α=0.10$) and retains substantial coverage. Compared to vanilla baselines, \textsc{Scope} accepts up to $2.4\times$ more judgments under the same risk constraint, demonstrating that BPE enables reliable and high-coverage LLM-based evaluation.

📄 PDF Abstract BibTeX arXiv:2602.13110

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

coverforest: Conformal Predictions with Random Forest in Python

2025-01-24 · Panisara Meehinkong, Donlapark Ponnoprat

Conformal prediction provides a framework for uncertainty quantification, specifically in the forms of prediction intervals and sets with distribution-free guaranteed coverage. While recent cross-conformal techniques suc…

Conformal PredictionPredictionPrediction IntervalsUncertainty Quantification

A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

2026-06-07 · Xiaoli Yu, Jiamiao Liu arxiv

Selective predictors answer on confident inputs and abstain elsewhere; deploying one safely needs a single finite-sample certificate that simultaneously upper-bounds the selected risk, lower-bounds the acceptance probabi…

Selective Conformal Risk Control

2025-12-14 · Yunpeng Xu, Wenge Guo, Zhi Wei arxiv

Reliable uncertainty quantification is essential for deploying machine learning systems in high-stakes domains. Conformal prediction provides distribution-free coverage guarantees but often produces overly large predicti…

CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging

2026-05-19 · Zhenyu Li, Aleksandar Cvejic, Zehui Chen, Peter Wonka arxiv

Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While rubric-based LLM judges improve interpre…

SCOPE: Sequential Conformal Probing for Reliable OOD Rejection in LLM Services

2026-06-19 · Zhuoyun Li, Boxuan Wang, Changshun Wu, Xiaowei Huang 외 arxiv

Rejecting inputs outside the defined in-distribution (IND) service scope is critical for large language model (LLM) services, where unsupported requests should be filtered before full generation. Existing out-of-distribu…