paper-with-me

Papers

Learning to Trust: How Humans Mentally Recalibrate AI Confidence Signals

2026-03-23 · ZhaoBin Li, Mark Steyvers arxiv

Productive human-AI collaboration requires appropriate reliance, yet contemporary AI systems are often miscalibrated, exhibiting systematic overconfidence or underconfidence. We investigate whether humans can learn to mentally recalibrate AI confidence signals through repeated experience. In a behavioral experiment (N = 200), participants predicted the AI's correctness across four AI calibration conditions: standard, overconfidence, underconfidence, and a counterintuitive "reverse confidence" mapping. Results demonstrate robust learning across all conditions, with participants significantly improving their accuracy, discrimination, and calibration alignment over 50 trials. We present a computational model utilizing a linear-in-log-odds (LLO) transformation and a Rescorla-Wagner learning rule to explain these dynamics. The model reveals that humans adapt by updating their baseline trust and confidence sensitivity, using asymmetric learning rates to prioritize the most informative errors. While humans can compensate for monotonic miscalibration, we identify a significant boundary in the reverse confidence scenario, where a substantial proportion of participants struggled to override initial inductive biases. These findings provide a mechanistic account of how humans adapt their trust in AI confidence signals through experience.

📄 PDF Abstract BibTeX arXiv:2603.22634

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making

2023-01-14 · Shuai Ma, Ying Lei, Xinru Wang, Chengbo Zheng 외

In AI-assisted decision-making, it is critical for human decision-makers to know when to trust AI and when to trust themselves. However, prior studies calibrated human trust only based on AI confidence indicating AI's co…

Decision Making

Improving Model Understanding and Trust with Counterfactual Explanations of Model Confidence

2022-06-06 · Thao Le, Tim Miller, Ronal Singh, Liz Sonenberg

In this paper, we show that counterfactual explanations of confidence scores help users better understand and better trust an AI model's prediction in human-subject studies. Showing confidence scores in human-agent inter…

counterfactualCounterfactual Explanationmodel

MACEst: The reliable and trustworthy Model Agnostic Confidence Estimator

2021-09-02 · Rhys Green, Matthew Rowe, Alberto Polleri

Reliable Confidence Estimates are hugely important for any machine learning model to be truly useful. In this paper, we argue that any confidence estimates based upon standard machine learning point prediction algorithms…

BIG-bench Machine Learning

Few-Shot Recalibration of Language Models

2024-03-27 · Xiang Lisa Li, Urvashi Khandelwal, Kelvin Guu

Recent work has uncovered promising ways to extract well-calibrated confidence estimates from language models (LMs), where the model's confidence score reflects how likely it is to be correct. However, while LMs may appe…

MathMMLU

Explaining Model Confidence Using Counterfactuals

2023-03-10 · Thao Le, Tim Miller, Ronal Singh, Liz Sonenberg

Displaying confidence scores in human-AI interaction has been shown to help build trust between humans and AI systems. However, most existing research uses only the confidence score as a form of communication. As confide…

counterfactualCounterfactual Explanationmodel