paper-with-me

홈 › Papers

Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness

2023-08-30 · Jiuhai Chen, Jonas Mueller

We introduce BSDetector, a method for detecting bad and speculative answers from a pretrained Large Language Model by estimating a numeric confidence score for any output it generated. Our uncertainty quantification technique works for any LLM accessible only via a black-box API, whose training data remains unknown. By expending a bit of extra computation, users of any LLM API can now get the same response as they would ordinarily, as well as a confidence estimate that cautions when not to trust this response. Experiments on both closed and open-form Question-Answer benchmarks reveal that BSDetector more accurately identifies incorrect LLM responses than alternative uncertainty estimation procedures (for both GPT-3 and ChatGPT). By sampling multiple responses from the LLM and considering the one with the highest confidence score, we can additionally obtain more accurate responses from the same LLM, without any extra training steps. In applications involving automated evaluation with LLMs, accounting for our confidence scores leads to more reliable evaluation in both human-in-the-loop and fully-automated settings (across both GPT 3.5 and 4).

📄 PDF Abstract BibTeX arXiv:2308.16175

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelUncertainty Quantification

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Discriminative Fine-Tuning Discriminative Fine-Tuning is a fine-tuning strategy that is used for ULMFiT type models. Instead of using the same learning rate…
GPT GPT is a Transformer-based architecture and training procedure for natural language processing tasks. Training follows a…
Residual Connection 설명 없음
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

2025-05-22 · Panagiotis Lymperopoulos, Vasanth Sarathy

Modern Large Language Models (LLMs) often require external tools, such as machine learning classifiers or knowledge retrieval systems, to provide accurate answers in domains where their pre-trained knowledge is insuffici…

Information RetrievalQuestion AnsweringRAGRetrieval+2

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

2024-10-04 · Robert E. Blackwell, Jon Barry, Anthony G. Cohn

Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, …

On the Role of Unobserved Sequences on Sample-based Uncertainty Quantification for LLMs

2025-10-06 · Lucie Kunitomo-Jacquin, Edison Marrese-Taylor, Ken Fukuda arxiv

Quantifying uncertainty in large language models (LLMs) is important for safety-critical applications because it helps spot incorrect answers, known as hallucinations. One major trend of uncertainty quantification method…

Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering

2025-09-18 · Yangyi Li, Mengdi Huai arxiv

Large language models (LLMs) have shown strong capabilities, enabling concise, context-aware answers in question answering (QA) tasks. The lack of transparency in complex LLMs has inspired extensive research aimed at dev…

Question Answering

Improving Uncertainty Quantification in Large Language Models via Semantic Embeddings

2024-10-30 · Yashvir S. Grewal, Edwin V. Bonilla, Thang D. Bui

Accurately quantifying uncertainty in large language models (LLMs) is crucial for their reliable deployment, especially in high-stakes applications. Current state-of-the-art methods for measuring semantic uncertainty in …

Question AnsweringUncertainty Quantification