paper-with-me

홈 › Papers

Bayesian Elicitation with LLMs: Model Size Helps, Extra "Reasoning" Doesn't Always

2026-04-02 · Luka Hobor, Mario Brcic, Mihael Kovac, Kristijan Poje arxiv

Large language models (LLMs) have been proposed as alternatives to human experts for estimating unknown quantities with associated uncertainty, a process known as Bayesian elicitation. We test this by asking eleven LLMs to estimate population statistics, such as health prevalence rates, personality trait distributions, and labor market figures, and to express their uncertainty as 95\% credible intervals. We vary each model's reasoning effort (low, medium, high) to test whether more "thinking" improves results. Our findings reveal three key results. First, larger, more capable models produce more accurate estimates, but increasing reasoning effort provides no consistent benefit. Second, all models are severely overconfident: their 95\% intervals contain the true value only 9--44\% of the time, far below the expected 95\%. Third, a statistical recalibration technique called conformal prediction can correct this overconfidence, expanding the intervals to achieve the intended coverage. In a preliminary experiment, giving models web search access degraded predictions for already-accurate models, while modestly improving predictions for weaker ones. Models performed well on commonly discussed topics but struggled with specialized health data. These results indicate that LLM uncertainty estimates require statistical correction before they can be used in decision-making.

📄 PDF Abstract BibTeX arXiv:2604.01896

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scalability of Bayesian Network Structure Elicitation with Large Language Models: a Novel Methodology and Comparative Analysis

2024-07-12 · Nikolay Babakov, Ehud Reiter, Alberto Bugarin

In this work, we propose a novel method for Bayesian Networks (BNs) structure elicitation that is based on the initialization of several LLMs with different experiences, independently querying them to create a structure …

Can LLMs Assist Expert Elicitation for Probabilistic Causal Modeling?

2025-04-14 · Olha Shaposhnyk, Daria Zahorska, Svetlana Yanushkevich

Objective: This study investigates the potential of Large Language Models (LLMs) as an alternative to human expert elicitation for extracting structured causal knowledge and facilitating causal modeling in biometric and …

Decision Making

Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation

2024-05-02 · David Eric Austin, Anton Korikov, Armin Toroghi, Scott Sanner

Designing preference elicitation (PE) methodologies that can quickly ascertain a user's top item preferences in a cold-start setting is a key challenge for building effective and personalized conversational recommendatio…

Bayesian OptimizationConversational RecommendationNatural Language InferenceThompson Sampling

On Bayesian Exponentially Embedded Family for Model Order Selection

2017-03-30 · Zhenghan Zhu, Steven Kay

In this paper, we derive a Bayesian model order selection rule by using the exponentially embedded family method, termed Bayesian EEF. Unlike many other Bayesian model selection methods, the Bayesian EEF can use vague pr…

Model Selection

Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen

2026-04-24 · Jon-Paul Cacioli arxiv

Verbal confidence elicitation is widely used to extract uncertainty estimates from LLMs. We tested whether seven instruction-tuned open-weight models (3-9B parameters, four families) produce verbalised confidence that me…