paper-with-me

Papers

When Confidence Signals Disagree: Local and Global Confidence in Autoregressive Language Models

2026-09-15 · Julio C. Amador Diaz Lopez arxiv

Modern predictive systems expose multiple quantities that are commonly interpreted as measures of confidence. However, these quantities can summarize different aspects of the predictive process. This distinction matters when confidence is used to evaluate reliability or inform downstream oversight and control. We investigate whether different confidence readouts are empirically interchangeable in an autoregressive language model by comparing local confidence, defined from the probability of the greedy-selected answer token, with global confidence, defined from modal-answer frequency under repeated sampling. Across MMLU and ARC Challenge, the two signals are weakly correlated and differ substantially in their association with correctness: global confidence is moderately associated with correctness, whereas local confidence shows little association. We further test whether question-level disagreement between the signals is associated with sampling instability. On ARC, larger local--global confidence gaps are associated with higher answer entropy, more distinct sampled answers, and lower modal-answer concentration. The gap--entropy association persists when disagreement and instability are estimated from disjoint stochastic samples, indicating that it is not explained by shared finite-sample variation. The corresponding relationship is substantially weaker on MMLU, where only 4% of questions exhibit sampling instability. These results show that confidence readouts derived from the same predictive system are not empirically interchangeable and that their disagreement can provide a diagnostic of unstable sampling behavior. Confidence should therefore be treated as an explicitly defined measurement rather than as a single intrinsic scalar property of a model, particularly when it is used to inform downstream evaluation, oversight, or control.

📄 PDF Abstract BibTeX arXiv:2609.16933

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Don't Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations

2024-03-06 · Abhishek Anand, Negar Mokhberian, Prathyusha Naresh Kumar, Anweasha Saha 외

Researchers have raised awareness about the harms of aggregating labels especially in subjective tasks that naturally contain disagreements among human annotators. In this work we show that models that are only provided …

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

2026-05-14 · Gaojie Jin, Yong Tao, Lijia Yu, Tianjin Huang arxiv

Jung et al. (2025) introduce a hypothesis testing framework for guaranteeing agreement between large language models (LLMs) and human judgments, relying on the assumption that the model's estimated confidence is monotoni…

Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA

2026-07-28 · Tom Saliencro, Rohan Desai, Priya Nair, Maya Lindqvist 외 arxiv

Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts $k$. Tokens differ in how uncertain the model is about them, so a single k over-spends on easy tokens and und…

Out-of-Distribution Detection

Localize-Then-Decide Guarantees for LLM Judgments

2026-08-26 · Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu 외 arxiv

Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work int…

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification

2026-05-19 · Rohit Sinha, Kunal Tilaganji, Tanuja Ganu, Nagarajan Natarajan 외 arxiv

Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Learned critics need extensive labeled d…