paper-with-me

홈 › Papers

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

2026-06-18 · Jiayi Wang, Xu-Yao Zhang arxiv

Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.

📄 PDF Abstract BibTeX arXiv:2606.19868

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting Uncertainty Estimation and Calibration of Large Language Models

2025-05-29 · Linwei Tao, Yi-Fan Yeh, Minjing Dong, Tao Huang 외

As large language models (LLMs) are increasingly deployed in high-stakes applications, robust uncertainty estimation is essential for ensuring the safe and trustworthy deployment of LLMs. We present the most comprehensiv…

Mixture-of-ExpertsMMLUQuantization

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

2026-01-06 · David Hartmann, Lena Pohlmann, Lelia Hanslik, Noah Gießing 외 arxiv

Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We concep…

Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models

2025-10-23 · Christian Hobelsberger, Theresa Winner, Andreas Nawroth, Oliver Mitevski 외 arxiv

Large language models (LLMs) produce outputs with varying levels of uncertainty, and, just as often, varying levels of correctness; making their practical reliability far from guaranteed. To quantify this uncertainty, we…

Efficient Non-Parametric Uncertainty Quantification for Black-Box Large Language Models and Decision Planning

2024-02-01 · Yao-Hung Hubert Tsai, Walter Talbott, Jian Zhang

Step-by-step decision planning with large language models (LLMs) is gaining attention in AI agent development. This paper focuses on decision planning with uncertainty estimation to address the hallucination problem in l…

AI AgentDecision MakingHallucinationUncertainty Quantification

pyPESTO: A modular and scalable tool for parameter estimation for dynamic models

2023-05-02 · Yannik Schälte, Fabian Fröhlich, Paul J. Jost, Jakob Vanhoefer 외

Mechanistic models are important tools to describe and understand biological processes. However, they typically rely on unknown parameters, the estimation of which can be challenging for large and complex systems. We pre…

parameter estimationUncertainty Quantification