paper-with-me

홈 › Papers

Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence?

2026-05-27 · Gabrielle Kaili-May Liu, Arman Cohan arxiv

LLMs' linguistically expressed confidence should faithfully reflect their intrinsic uncertainty. While recent work shows LLMs struggle to use epistemic markers (e.g., "it is likely...") in a human-aligned fashion, it remains unclear whether models can apply their own linguistic confidence framework to associate markers with specific confidence levels in a stable and generalizable way, and how contextual features impact this ability. We conduct the first systematic study of this question, formalizing _marker internal confidence_ (MIC) as the estimated intrinsic confidence a model associates with a specific epistemic marker in a given task domain. We present 7 metrics to evaluate the stability of MICs within and across distributions. Applying our analysis framework to diverse models and tasks, we find that LLMs remain faithfully miscalibrated even under model-centric interpretation of marker meanings, struggling to differentiate markers by internal confidence across distributions despite preserving a somewhat consistent ranking order across tasks. This supplies critical, complementary evidence to existing work toward a holistic understanding of faithful calibration in LLMs, emphasizing the need for more aligned and stable marker use to improve trustworthiness and reliability.

📄 PDF Abstract BibTeX arXiv:2605.28778

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?

2025-05-30 · Jiayu Liu, Qing Zong, Weiqi Wang, Yangqiu Song

As large language models (LLMs) are increasingly used in high-stakes domains, accurately assessing their confidence is crucial. Humans typically express confidence through epistemic markers (e.g., "fairly confident") ins…

Question Answering

Humans overrely on overconfident language models, across languages

2025-07-08 · Neil Rathi, Dan Jurafsky, Kaitlyn Zhou arxiv

As large language models (LLMs) are deployed globally, it is crucial that their responses are calibrated across languages to accurately convey uncertainty and limitations. Prior work shows that LLMs are linguistically ov…

Secret Keepers: The Impact of LLMs on Linguistic Markers of Personal Traits

2024-03-30 · Zhivar Sourati, Meltem Ozcan, Colin McDaniel, Alireza Ziabari 외

Prior research has established associations between individuals' language usage and their personal traits; our linguistic patterns reveal information about our personalities, emotional states, and beliefs. However, with …

Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with Dementia

2023-10-15 · Dimitris Gkoumas, Matthew Purver, Maria Liakata

Dementia is associated with language disorders which impede communication. Here, we automatically learn linguistic disorder patterns by making use of a moderately-sized pre-trained language model and forcing it to focus …

Language ModelingLanguage Modelling

Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs

2026-03-13 · Gustavo Lúcius Fernandes, Jeiverson C. V. M. Santos, Pedro O. S. Vaz-de-Melo arxiv

Large language models (LLMs) are increasingly used to assess moral or ethical statements, yet their judgments may reflect social and linguistic biases. This work presents a controlled, sentence-level study of how grammat…