paper-with-me

Papers

Localize-Then-Decide Guarantees for LLM Judgments

2026-08-26 · Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin arxiv

Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence-thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans. However, this assumption can break down when the number of candidate responses increases, since distributing probability mass across many alternatives can distort confidence estimates. To address this issue, we propose a Localize-Then-Decide framework. First, conformal prediction localizes a small shortlist that contains the human-preferred response with high probability. Then, a calibrated confidence-based rule selectively chooses a single response from this shortlist or abstains. This design restores the monotonic relationship between confidence and disagreement risk and enables high-probability agreement guarantees. Experiments with multiple candidate sizes across several datasets and judge LLMs demonstrate that our framework consistently achieves higher guarantee success rates and substantially higher coverage than single-stage baselines.

📄 PDF Abstract BibTeX arXiv:2608.25824

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Combining Experts' Causal Judgments

2020-05-20 · Dalal Alrajeh, Hana Chockler, Joseph Y. Halpern

Consider a policymaker who wants to decide which intervention to perform in order to change a currently undesirable situation. The policymaker has at her disposal a team of experts, each with their own understanding of t…

Collecting Consistently High Quality Object Tracks with Minimal Human Involvement by Using Self-Supervised Learning to Detect Tracker Errors

2024-05-06 · Samreen Anjum, Suyog Jain, Danna Gurari

We propose a hybrid framework for consistently producing high-quality object tracks by combining an automated object tracker with little human input. The key idea is to tailor a module for each dataset to intelligently d…

ObjectSelf-Supervised Learning

Street-Level AI: Are Large Language Models Ready for Real-World Judgments?

2025-08-11 · Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das arxiv

A surge of recent work explores the ethical and societal implications of large-scale AI models that make "moral" judgments. Much of this literature focuses either on alignment with human judgments through various thought…

Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction

2026-08-06 · Anton Conrad, Rustam Isaev, Denis Belomestny, Eric Moulines 외 arxiv

Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free c…

Studying Summarization Evaluation Metrics in the Appropriate Scoring Range

2019-07-01 · ACL 2019 7 · Maxime Peyrard

In summarization, automatic evaluation metrics are usually compared based on their ability to correlate with human judgments. Unfortunately, the few existing human judgment datasets have been created as by-products of th…