paper-with-me

홈 › Papers

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

2026-07-07 · Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico arxiv

Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding. Second, prompting models to reason in English closes the UE performance gap between low and high-resource languages, demonstrating that generation language matters more than the question language. Third, the choice of UE method should depend on model scale: at smaller scales, open-box probability-based methods outperform alternatives; at larger scales, closed-box self-verbalized uncertainty becomes superior. Finally, we provide an analysis of threshold selection for selective prediction, offering guidance on calibrating abstention in multilingual settings.

📄 PDF Abstract BibTeX arXiv:2607.06327

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Uncertainty Estimation Scales with Sampling in Reasoning Models

2026-03-19 · Maksym Del, Markus Kängsepp, Marharyta Domnich, Ardi Tampuu 외 arxiv

Uncertainty estimation is critical for deploying reasoning language models, yet remains poorly understood under extended chain-of-thought reasoning. We study parallel sampling as a fully black-box approach using verbaliz…

Revisiting Uncertainty Estimation and Calibration of Large Language Models

2025-05-29 · Linwei Tao, Yi-Fan Yeh, Minjing Dong, Tao Huang 외

As large language models (LLMs) are increasingly deployed in high-stakes applications, robust uncertainty estimation is essential for ensuring the safe and trustworthy deployment of LLMs. We present the most comprehensiv…

Mixture-of-ExpertsMMLUQuantization

Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

2025-03-20 · Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen 외

Large Language Models (LLMs) excel in text generation, reasoning, and decision-making, enabling their adoption in high-stakes domains such as healthcare, law, and transportation. However, their reliability is a major con…

Computational EfficiencyDecision MakingText GenerationUncertainty Quantification

Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation

2026-08-20 · Eric Bigelow, Amir Zur, Satchel Grant, Tal Haklay 외 arxiv

LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produce for a given question, i.e., its uncertainty. Resampling-based analyses characte…

Text Generation

What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study

2026-07-07 · Robert Richardson arxiv

Evaluating uncertainty in AI-generated SQL queries requires estimating whether a query is correct, where correct means it executes to the same result as a human-written reference. We study which signals predict correctne…