paper-with-me

Papers

Enhancing Healthcare LLM Trust with Atypical Presentations Recalibration

2024-09-05 · Jeremy Qin, Bang Liu, Quoc Dinh Nguyen

Black-box large language models (LLMs) are increasingly deployed in various environments, making it essential for these models to effectively convey their confidence and uncertainty, especially in high-stakes settings. However, these models often exhibit overconfidence, leading to potential risks and misjudgments. Existing techniques for eliciting and calibrating LLM confidence have primarily focused on general reasoning datasets, yielding only modest improvements. Accurate calibration is crucial for informed decision-making and preventing adverse outcomes but remains challenging due to the complexity and variability of tasks these models perform. In this work, we investigate the miscalibration behavior of black-box LLMs within the healthcare setting. We propose a novel method, \textit{Atypical Presentations Recalibration}, which leverages atypical presentations to adjust the model's confidence estimates. Our approach significantly improves calibration, reducing calibration errors by approximately 60\% on three medical question answering datasets and outperforming existing methods such as vanilla verbalized confidence, CoT verbalized confidence and others. Additionally, we provide an in-depth analysis of the role of atypicality within the recalibration framework.

📄 PDF Abstract BibTeX arXiv:2409.03225

Code (1)

jeremy-qin/medical_confidence_elicitation 공식 구현 pytorch

Tasks

Decision MakingMedical Question AnsweringQuestion Answering

Similar Papers 제목 키워드 기반

Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO

2024-11-01 · Macarious Hui, Jinda Zhang, Aanchan Mohan

Individuals with cerebral palsy (CP) and amyotrophic lateral sclerosis (ALS) frequently face challenges with articulation, leading to dysarthria and resulting in atypical speech patterns. In healthcare settings, communic…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModellingLarge Language Model+2

Aligning Large Language Models with Healthcare Stakeholders: A Pathway to Trustworthy AI Integration

2025-05-02 · Kexin Ding, Mu Zhou, Akshay Chaudhari, Shaoting Zhang 외

The wide exploration of large language models (LLMs) raises the awareness of alignment between healthcare stakeholder preferences and model outputs. This alignment becomes a crucial foundation to empower the healthcare w…

From Uncertainty to Precision: Enhancing Binary Classifier Performance through Calibration

2024-02-12 · Agathe Fernandes Machado, Arthur Charpentier, Emmanuel Flachaire, Ewen Gallic 외

The assessment of binary classifier performance traditionally centers on discriminative ability using metrics, such as accuracy. However, these metrics often disregard the model's inherent uncertainty, especially when de…

Decision Making

Trustworthy Machine Learning via Memorization and the Granular Long-Tail: A Survey on Interactions, Tradeoffs, and Beyond

2025-03-10 · Qiongxiu Li, Xiaoyu Luo, Yiyi Chen, Johannes Bjerva

The role of memorization in machine learning (ML) has garnered significant attention, particularly as modern models are empirically observed to memorize fragments of training data. Previous theoretical analyses, such as …

AttributeFairnessMemorization

Uncertainty Reliability Under Domain Shift: An Investigation for Data-Driven Blood Pressure Estimation in Photoplethysmography

2026-05-18 · Mohammad Moulaeifard, Ciaran Bench, Philip J. Aston, Nils Strodthoff arxiv

Uncertainty quantification (UQ) is critical for safety-critical domains like healthcare, yet it is rarely evaluated under realistic out-of-distribution (OOD) conditions. Here, we assessed predictive performance and uncer…

Blood pressure estimation