paper-with-me

홈 › Papers

Human-Alignment, Calibration, and Activation Patterns in Large Language Model Uncertainty

2026-05-29 · Kyle Moore, Jesse Roberts, Daryl Watson, William Ward, Grayson Heyboer arxiv

Uncertainty Quantification is a large and growing subfield of large language model behavioral analysis. Primarily to recognize and combat hallucination, the field has largely focused on measuring and improving calibration, the accuracy of uncertainty judgments to task efficacy. In this work, we investigate the relatively underexplored question of how similar large language model uncertainty is to human uncertainty. We investigate the presence and strength of human-similar uncertainty signals, deemed uncertainty alignment, in large language model overt behavior and internal activation patterns. We identify whether the models show evidence of simultaneous alignment and calibration on a variety of datasets covering both multiple choice and open ended factual recall. And we characterize the effect of instruct fine-tuning on each of these facets.

📄 PDF Abstract BibTeX arXiv:2605.30675

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning

2025-09-26 · Aayush Mishra, Daniel Khashabi, Anqi Liu arxiv

Supervised Fine-Tuning (SFT) is used to specialize model behavior by training weights to produce intended target responses for queries. In contrast, In-Context Learning (ICL) adapts models during inference with instructi…

LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulations

2026-03-02 · Veronika Solopova, Viktoria Skorik, Maksym Tereshchenko, Alina Haidun 외 arxiv

Large language models (LLMs) are increasingly proposed as agents in strategic decision environments, yet their behavior in structured geopolitical simulations remains under-researched. We evaluate six popular state-of-th…

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

2026-07-28 · Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson arxiv

Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too acc…

Mirror-Neuron Patterns in AI Alignment

2025-10-23 · Robyn Wyrick arxiv

As artificial intelligence (AI) advances toward superhuman capabilities, aligning these systems with human values becomes increasingly critical. Current alignment strategies rely largely on externally specified constrain…

Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time Steering

2026-02-05 · Miranda Muqing Miao, Young-Min Cho, Lyle Ungar arxiv

Large language models (LLMs) exhibit persistent miscalibration, especially after instruction tuning and preference alignment. Modified training objectives can improve calibration, but retraining is expensive. Inference-t…