paper-with-me

홈 › Papers

Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control

2025-08-07 · Yuanchang Ye arxiv

This study introduces a significance testing-enhanced conformal prediction (CP) framework to improve trustworthiness of large language models (LLMs) in multiple-choice question answering (MCQA). While LLMs have been increasingly deployed in disciplinary QA scenarios, hallucination and nonfactual generation substantially compromise response reliability. Although CP provides statistically rigorous marginal coverage guarantees for prediction sets, and significance testing offers established statistical rigor, their synergistic integration remains unexplored. To mitigate hallucination and factual inaccuracies, our framework integrates $p$-value computation with conformity scoring through self-consistency resampling of MCQA responses. This approach calculates option frequencies to address LLMs' black-box nature, subsequently constructing prediction sets via null hypothesis testing ($\mathcal{H}_0$) with empirically derived $p$-values. Evaluations on MMLU and MMLU-Pro benchmarks using off-the-shelf LLMs demonstrate: (1) The enhanced CP achieves user-specified empirical miscoverage rates; (2) Test-set average prediction set size (APSS) decreases monotonically with increasing risk levels ($α$), validating APSS as an effective uncertainty metric. This work establishes a principled statistical framework for trustworthy LLM deployment in high-stakes QA applications.

📄 PDF Abstract BibTeX arXiv:2508.10022

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Conformal Prediction with Large Language Models for Multi-Choice Question Answering

2023-05-28 · Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu 외

As large language models continue to be widely developed, robust uncertainty quantification techniques will become crucial for their safe deployment in high-stakes scenarios. In this work, we explore how conformal predic…

Conformal PredictionMultiple-choicePredictionQuestion Answering+1

Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework

2025-03-07 · Yusong Ke, Hongru Lin, Yuting Ruan, Junya Tang 외

Large language models (LLMs) are increasingly adopted in medical question-answering (QA) scenarios. However, LLMs can generate hallucinations and nonfactual information, undermining their trustworthiness in high-stakes m…

Conformal PredictionMedical Question AnsweringMedQAMMLU+4

Conformal Sets in Multiple-Choice Question Answering under Black-Box Settings with Provable Coverage Guarantees

2025-08-07 · Guang Yang, Xinyang Liu arxiv

Large Language Models (LLMs) have shown remarkable progress in multiple-choice question answering (MCQA), but their inherent unreliability, such as hallucination and overconfidence, limits their application in high-risk …

Question Answering

Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models

2026-04-15 · Aleksandr Rubashevskii, Dzianis Piatrashyn, Preslav Nakov, Maxim Panov arxiv

Large language models (LLMs) are prone to generating factually incorrect outputs. Recent work has applied conformal prediction to provide uncertainty estimates and statistical guarantees for the factuality of LLM generat…

Question Answering

Monty Hall and Optimized Conformal Prediction to Improve Decision-Making with LLMs

2024-12-31 · Harit Vishwakarma, Alan Mishler, Thomas Cook, Niccolò Dalmasso 외

Large language models (LLMs) are empowering decision-making in several applications, including tool or API usage and answering multiple-choice questions (MCQs). However, they often make overconfident, incorrect predictio…

Conformal PredictionDecision MakingMMLUMultiple-choice+3