paper-with-me

홈 › Papers

How Confident Is the First Token? An Uncertainty-Calibrated Prompt Optimization Framework for Large Language Model Classification and Understanding

2026-02-23 · Wei Chen, Guoyang Ju, Yuanyuan Qi arxiv

With the widespread adoption of large language models (LLMs) in natural language processing, prompt engineering and retrieval-augmented generation (RAG) have become mainstream to enhance LLMs' performance on complex tasks. However, LLMs generate outputs autoregressively, leading to inevitable output uncertainty. Since model performance is highly sensitive to prompt design, precise uncertainty measurement is crucial for reliable prompt optimization. For multi-class multiple-choice (understanding) tasks, conventional uncertainty measures (e.g., entropy) based on output probabilities treat all classes equally and ignore class prior differences in pretraining corpora. This failure to distinguish spurious confidence (from priors) from true certainty (from contextual understanding) results in poor confidence calibration. To address this, we propose Log-Scale Focal Uncertainty (LSFU), a first-token-based metric inspired by focal loss. LSFU incorporates label prior probabilities as a risk-modulation factor to suppress noise from high-frequency classes and emphasize risk for low-frequency long-tail classes, with a dynamic weighting mechanism unifying the measurement scale. Based on LSFU, we further propose the uncertainty-calibrated prompt optimization framework (UCPOF), which leverages the first token of model outputs to select high-quality exemplars and dynamically optimize prompts. Comprehensive evaluations show UCPOF improves average accuracy by 6.03% over few-shot baselines, surpasses always-on full RAG by 5.75% in overall average accuracy, and reduces the average retrieval trigger rate by 50.66%. By adaptively triggering RAG only for high-uncertainty samples, our framework significantly lowers computational costs while maintaining state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2603.18009

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Evaluating language models as risk scores

2024-07-19 · André F. Cruz, Moritz Hardt, Celestine Mendler-Dünner

Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks. Conditioned on a question and answer-key, does the most likely token match the ground truth? Such benchmarks necessari…

Multiple-choiceQuestion Answering

DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metrics

2024-07-08 · Luke Yoffe, Alfonso Amayuelas, William Yang Wang

Multi-agent debates have been introduced to improve the accuracy of Large Language Models (LLMs) by having multiple agents discuss solutions to a problem over several rounds of debate. However, models often generate inco…

Language ModelingLanguage ModellingLarge Language Model

Overconfident and Unconfident AI Hinder Human-AI Collaboration

2024-02-12 · Jingshu Li, Yitian Yang, Renwen Zhang, Yi-chieh Lee

AI transparency is a central pillar of responsible AI deployment and effective human-AI collaboration. A critical approach is communicating uncertainty, such as displaying AI's confidence level, or its correctness likeli…

Calibrated Reliable Regression using Maximum Mean Discrepancy

2020-06-18 · NeurIPS 2020 12 · Peng Cui, Wen-Bo Hu, Jun Zhu

Accurate quantification of uncertainty is crucial for real-world applications of machine learning. However, modern deep neural networks still produce unreliable predictive uncertainty, often yielding over-confident predi…

BIG-bench Machine LearningPrediction Intervalsregression

On Verbalized Confidence Scores for LLMs

2024-12-19 · Daniel Yang, Yao-Hung Hubert Tsai, Makoto Yamada

The rise of large language models (LLMs) and their tight integration into our daily life make it essential to dedicate efforts towards their trustworthiness. Uncertainty quantification for LLMs can establish more human t…

Uncertainty Quantification