paper-with-me

홈 › Papers

Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation

2025-10-02 · Mykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp Hochreiter arxiv

Hallucinations are a common issue that undermine the reliability of large language models (LLMs). Recent studies have identified a specific subset of hallucinations, known as confabulations, which arise due to predictive uncertainty of LLMs. To detect confabulations, various methods for estimating predictive uncertainty in natural language generation (NLG) have been developed. These methods are typically evaluated by correlating uncertainty estimates with the correctness of generated text, with question-answering (QA) datasets serving as the standard benchmark. However, commonly used approximate correctness functions have substantial disagreement between each other and, consequently, in the ranking of the uncertainty estimation methods. This allows one to inflate the apparent performance of uncertainty estimation methods. We propose using several alternative risk indicators for risk correlation experiments that improve robustness of empirical assessment of UE algorithms for NLG. For QA tasks, we show that marginalizing over multiple LLM-as-a-judge variants leads to reducing the evaluation biases. Furthermore, we explore structured tasks as well as out of distribution and perturbation detection tasks which provide robust and controllable risk indicators. Finally, we propose to use an Elo rating of uncertainty estimation methods to give an objective summarization over extensive evaluation settings.

📄 PDF Abstract BibTeX arXiv:2510.02279

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning

2020-02-15 · ICLR 2020 1 · Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry Vetrov

Uncertainty estimation and ensembling methods go hand-in-hand. Uncertainty estimation is one of the main benchmarks for assessment of ensembling performance. At the same time, deep learning ensembles have provided state-…

image-classificationImage Classification

Towards Harmonized Uncertainty Estimation for Large Language Models

2025-05-25 · Rui Li, Jing Long, Muge Qi, Heming Xia 외

To facilitate robust and trustworthy deployment of large language models (LLMs), it is essential to quantify the reliability of their generations through uncertainty estimation. While recent efforts have made significant…

IoT Device Identification with Machine Learning: Common Pitfalls and Best Practices

2026-01-28 · Kahraman Kostas, Rabia Yasa Kostas arxiv

This paper critically examines the device identification process using machine learning, addressing common pitfalls in existing literature. We analyze the trade-offs between identification methods (unique vs. class based…

Data Augmentation

Pitfalls of topology-aware image segmentation

2024-12-19 · Alexander H. Berger, Laurin Lux, Alexander Weers, Martin Menten 외

Topological correctness, i.e., the preservation of structural integrity and specific characteristics of shape, is a fundamental requirement for medical imaging tasks, such as neuron or vessel segmentation. Despite the re…

BenchmarkingImage SegmentationMedical Image SegmentationSegmentation+1

ValUES: A Framework for Systematic Validation of Uncertainty Estimation in Semantic Segmentation

2024-01-16 · Kim-Celine Kahl, Carsten T. Lüth, Maximilian Zenk, Klaus Maier-Hein 외

Uncertainty estimation is an essential and heavily-studied component for the reliable application of semantic segmentation methods. While various studies exist claiming methodological advances on the one hand, and succes…

Active LearningSemantic Segmentation