paper-with-me

Papers

Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs

2026-06-21 · Weihong Qi, Kristina Lerman arxiv

Large language models (LLMs) frequently encounter inputs that disagree with their prior outputs, through user pushback, retrieved documents, or web search results. While the way they resolve such conflicts -- a process we frame as cognitive dissonance resolution -- has been characterized behaviorally, its connection to internal model uncertainty is not well understood. To study this systematically, we vary persuasion attempts along two dimensions, source authority and evidence quality, across 12 health-science claims of stratified epistemic status. Dissonance can be resolved through persuasion, backfire, or immunity. We introduce Trust Elasticity (TE), an econometrics-inspired measure of how readily a model is persuaded toward conflicting evidence. Across four LLMs, TE varies substantially, while clearly false claims elicit near-zero TE across all models. On two open-weight models, we further find that this variation is associated with two complementary internal uncertainty indicators, Confidence Miscalibration in Qwen and Internal Uncertainty Change in Llama. These results link cross-model behavioral variation to a measurable internal property and point to interventions targeting internal uncertainty as future work.

📄 PDF Abstract BibTeX arXiv:2606.22633

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Cognitive Circuit Breaker: A Systems Engineering Framework for Intrinsic AI Reliability

2026-04-15 · Jonathan Pan arxiv

As Large Language Models (LLMs) are increasingly deployed in mission-critical software systems, detecting hallucinations and ``faked truthfulness'' has become a paramount engineering challenge. Current reliability archit…

Impact of different belief facets on agents' decision -- a refined cognitive architecture to model the interaction between organisations' institutional characteristics and agents' behaviour

2020-04-24 · Amir Hosein Afshar Sedigh, Martin K. Purvis, Bastin Tony Roy Savarimuthu, Christopher K. Frantz 외

This paper presents a conceptual refinement of agent cognitive architecture inspired from the beliefs-desires-intentions (BDI) and the theory of planned behaviour (TPB) models, with an emphasis on different belief facets…

Fairness

Characterizing Social Imaginaries and Self-Disclosures of Dissonance in Online Conspiracy Discussion Communities

2021-07-21 · Shruti Phadke, Mattia Samory, Tanushree Mitra

Online discussion platforms offer a forum to strengthen and propagate belief in misinformed conspiracy theories. Yet, they also offer avenues for conspiracy theorists to express their doubts and experiences of cognitive …

2k

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

2026-02-21 · Zhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi 외 arxiv

Active Learning (AL) reduces annotation costs in medical imaging by selecting only the most informative samples for labeling, but suffers from cold-start when labeled data are scarce. Vision-Language Models (VLMs) addres…

Active Learning

Uncertainty-Aware Reliable Text Classification

2021-07-15 · Yibo Hu, Latifur Khan

Deep neural networks have significantly contributed to the success in predictive accuracy for classification tasks. However, they tend to make over-confident predictions in real-world settings, where domain shifting and …

ClassificationOutlier DetectionOut of Distribution (OOD) Detectiontext-classification+1