paper-with-me

홈 › Papers

Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework

2025-03-07 · Yusong Ke, Hongru Lin, Yuting Ruan, Junya Tang, Li Li

Large language models (LLMs) are increasingly adopted in medical question-answering (QA) scenarios. However, LLMs can generate hallucinations and nonfactual information, undermining their trustworthiness in high-stakes medical tasks. Conformal Prediction (CP) provides a statistically rigorous framework for marginal (average) coverage guarantees but has limited exploration in medical QA. This paper proposes an enhanced CP framework for medical multiple-choice question-answering (MCQA) tasks. By associating the non-conformance score with the frequency score of correct options and leveraging self-consistency, the framework addresses internal model opacity and incorporates a risk control strategy with a monotonic loss function. Evaluated on MedMCQA, MedQA, and MMLU datasets using four off-the-shelf LLMs, the proposed method meets specified error rate guarantees while reducing average prediction set size with increased risk level, offering a promising uncertainty evaluation metric for LLMs.

📄 PDF Abstract BibTeX arXiv:2503.05505

Code (0)

등록된 구현이 없습니다.

Tasks

Conformal PredictionMedical Question AnsweringMedQAMMLUMultiple-choiceMultiple Choice Question Answering (MCQA)PredictionQuestion Answering

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

2026-03-02 · Xiang Zheng, Han Li, Wenjie Luo, Weiqi Zhai 외 arxiv

Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We introduce \textsc{ClinConsensus}, a Chine…

Generation-Time vs. Post-hoc Citation: A Holistic Evaluation of LLM Attribution

2025-09-25 · Yash Saxena, Raviteja Bommireddy, Ankur Padia, Manas Gaur arxiv

Trustworthy Large Language Models (LLMs) must cite human-verifiable sources in high-stakes domains such as healthcare, law, academia, and finance, where even small errors can have severe consequences. Practitioners and r…

MediX-R1: Open Ended Medical Reinforcement Learning

2026-02-26 · Sahal Shaji Mullappilly, Mohammed Irfan Kurpath, Omair Mohamed, Mohamed Zidan 외 arxiv

We introduce MediX-R1, an open-ended Reinforcement Learning (RL) framework for medical multimodal large language models (MLLMs) that enables clinically grounded, free-form answers beyond multiple-choice formats. MediX-R1…

Reinforcement Learning

Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation

2025-01-12 · Shunfan Zheng, Xiechi Zhang, Gerard de Melo, Xiaoling Wang 외

In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount. Existing benchmarks often focus on fixe…

AttributeMultiple-choice

MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation

2025-12-18 · Pengyu Wang, Shuchang Ye, Usman Naseem, Jinman Kim arxiv

Medical report generation aims to automatically produce radiology-style reports from medical images, supporting efficient and accurate clinical decision-making.However, existing approaches predominately rely on token-lev…

Medical Report GenerationReinforcement Learning