paper-with-me

홈 › Papers

Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning

2026-09-15 · Yuwei Liang, Jian Liang, Dapeng Hu, Yinuo Xu, Ran He arxiv

Test-time prompt tuning (TPT) enables adaptation on a single test instance, achieving improved accuracy but often sacrificing calibration performance. Most existing calibration methods introduce additional regularization terms to promote dispersion across text embeddings and reduce calibration error, yet these methods often suffer from a drop in accuracy. Motivated by the well-calibrated nature of zero-shot predictions, we propose CoTS, a simple yet effective post-hoc calibration method that preserves accuracy. Specifically, CoTS applies temperature scaling to minimize the confidence gap between adapted and zero-shot predictions. To fully exploit the potential of multiple augmentations during adaptation, we introduce a weak-strong ensemble strategy that further boosts accuracy. We then apply CoTS to this ensemble, termed E-CoTS, to maintain its well-calibrated property. Extensive experiments on diverse datasets and backbones show that our approaches effectively mitigate miscalibration without compromising primary accuracy. For instance, E-CoTS reduces the average expected calibration error of TPT from 11.90% to 5.38% on ImageNet variants, while even increasing accuracy from 60.74% to 62.95%. Moreover, when integrated with existing calibration methods, E-CoTS usually enhances both accuracy and calibration simultaneously.

📄 PDF Abstract BibTeX arXiv:2609.17386

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Domain-adaptive and Subgroup-specific Cascaded Temperature Regression for Out-of-distribution Calibration

2024-02-14 · Jiexin Wang, Jiahao Chen, Bing Su

Although deep neural networks yield high classification accuracy given sufficient training data, their predictions are typically overconfident or under-confident, i.e., the prediction confidences cannot truly reflect the…

Data Augmentationregression

Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator

2025-05-22 · Beier Luo, Shuoyuan Wang, Yixuan Li, Hongxin Wei

Post-training of large language models is essential for adapting pre-trained language models (PLMs) to align with human preferences and downstream tasks. While PLMs typically exhibit well-calibrated confidence, post-trai…

Reducing Overconfident Errors outside the Known Distribution

2019-05-01 · ICLR 2019 5 · Zhizhong Li, Derek Hoiem

Intuitively, unfamiliarity should lead to lack of confidence. In reality, current algorithms often make highly confident yet wrong predictions when faced with unexpected test samples from an unknown distribution differen…

Domain AdaptationNovelty Detection

Calibrating Language Models with Adaptive Temperature Scaling

2024-09-29 · Johnathan Xie, Annie S. Chen, Yoonho Lee, Eric Mitchell 외

The effectiveness of large language models (LLMs) is not only measured by their ability to generate accurate outputs but also by their calibration-how well their confidence scores reflect the probability of their outputs…

Unsupervised Pre-training

Attended Temperature Scaling: A Practical Approach for Calibrating Deep Neural Networks

2018-10-27 · Azadeh Sadat Mozafari, Hugo Siqueira Gomes, Wilson Leão, Steeven Janny 외

Recently, Deep Neural Networks (DNNs) have been achieving impressive results on wide range of tasks. However, they suffer from being well-calibrated. In decision-making applications, such as autonomous driving or medical…

Autonomous DrivingDecision MakingLesion Detection