paper-with-me

홈 › Papers

When to Trust Confidence Thresholding: Calibration Diagnostics for Pseudo-Labelled Regression

2026-05-12 · Marcell T. Kurbucz arxiv

Calibrated probability outputs of trained classifiers are increasingly used as inputs to downstream regression estimands such as effects, prevalences, or disparities for a latent group observed only on a small labelled subset. A standard practice is to threshold the calibrated score at a confidence cutoff and treat the hard label as the truth. Building on a recent identification result for the underlying moment equation, we develop a calibration-aware diagnostic apparatus for pseudo-labelling pipelines. We derive a closed-form expression for the attenuation bias that confidence thresholding induces in the downstream regression coefficient, and show that the bias can be predicted, before any inference is run, from the residual score variance $V^{*}=\mathbb{E}[\operatorname{Var}(p\mid X)]$ on the unlabelled set after partialling out the downstream controls $X$. We further obtain a sharp sensitivity bound under bounded calibration drift, and identify the boundary $V^{*}=0$, which holds iff $p$ is a deterministic function of $X$; this motivates a structural separation between classifier features $W$ and downstream controls $X\subsetneq W$. Five controlled simulations and a UCI Adult illustration trace the predictions. The contribution is operational: a $(V^{*}, κ)$ decision rule that practitioners can compute from any classifier output to decide whether confidence thresholding is safe.

📄 PDF Abstract BibTeX arXiv:2605.12780

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Orchestrator-Agent Trust: A Modular Agentic AI Visual Classification System with Trust-Aware Orchestration and RAG-Based Reasoning

2025-07-09 · Konstantinos I. Roumeliotis, Ranjan Sapkota, Manoj Karkee, Nikolaos D. Tselikas

Modern Artificial Intelligence (AI) increasingly relies on multi-agent architectures that blend visual and language understanding. Yet, a pressing challenge remains: How can we trust these agents especially in zero-shot …

BenchmarkingImage RetrievalOptical Character Recognition (OCR)RAG+3

The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

2026-06-05 · Nishant Subramani, Palash Goyal, Yiwen Song, Mani Malek 외 arxiv

As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is a good proxy for trust: well-calibrated confidence estimates help inform the risk…

Scientific Document SummarizationQuestion Answering

We Care Each Pixel: Calibrating on Medical Segmentation Model

2025-03-07 · Wenhao Liang, Wei zhang, Lin Yue, Miao Xu 외

Medical image segmentation is fundamental for computer-aided diagnostics, providing accurate delineation of anatomical structures and pathological regions. While common metrics such as Accuracy, DSC, IoU, and HD primaril…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models

2025-07-12 · Anita Kriz, Elizabeth Laura Janes, Xing Shen, Tal Arbel arxiv

Multimodal large language models (MLLMs) hold considerable promise for applications in healthcare. However, their deployment in safety-critical settings is hindered by two key limitations: (i) sensitivity to prompt desig…

Visual Question AnsweringZero-shot GeneralizationReinforcement LearningPrompt Engineering

Calibration Meets Explanation: A Simple and Effective Approach for Model Confidence Estimates

2022-11-06 · Dongfang Li, Baotian Hu, Qingcai Chen

Calibration strengthens the trustworthiness of black-box models by producing better accurate confidence estimates on given examples. However, little is known about if model explanations can help confidence calibration. I…