paper-with-me

홈 › Papers

How Uncertainty Estimation Scales with Sampling in Reasoning Models

2026-03-19 · Maksym Del, Markus Kängsepp, Marharyta Domnich, Ardi Tampuu, Lisa Yankovskaya, Meelis Kull, Mark Fishel arxiv

Uncertainty estimation is critical for deploying reasoning language models, yet remains poorly understood under extended chain-of-thought reasoning. We study parallel sampling as a fully black-box approach using verbalized confidence and self-consistency. Across three reasoning models and 17 tasks spanning mathematics, STEM, and humanities, we characterize how these signals scale. Both self-consistency and verbalized confidence scale in reasoning models, but self-consistency exhibits lower initial discrimination and lags behind verbalized confidence under moderate sampling. Most uncertainty gains, however, arise from signal combination: with just two samples, a hybrid estimator improves AUROC by up to $+12$ on average and already outperforms either signal alone even when scaled to much larger budgets, after which returns diminish. These effects are domain-dependent: in mathematics, the native domain of RLVR-style post-training, reasoning models achieve higher uncertainty quality and exhibit both stronger complementarity and faster scaling than in STEM or humanities.

📄 PDF Abstract BibTeX arXiv:2603.19118

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting Uncertainty Estimation and Calibration of Large Language Models

2025-05-29 · Linwei Tao, Yi-Fan Yeh, Minjing Dong, Tao Huang 외

As large language models (LLMs) are increasingly deployed in high-stakes applications, robust uncertainty estimation is essential for ensuring the safe and trustworthy deployment of LLMs. We present the most comprehensiv…

Mixture-of-ExpertsMMLUQuantization

SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio

2026-04-07 · Satwik Pandey, Suresh Raghu, Shashwat Pandey arxiv

Uncertainty estimation for reasoning language models remains difficult to deploy in practice: sampling-based methods are computationally expensive, while common single-pass proxies such as verbalized confidence or trace …

Multi-Robot Allocation for Information Gathering in Non-Uniform Spatiotemporal Environments

2025-09-26 · Kaleb Ben Naveed, Haejoon Lee, Dimitra Panagou arxiv

Autonomous robots are increasingly deployed to estimate spatiotemporal fields (e.g., wind, temperature, gas concentration) that vary across space and time. We consider environments divided into non-overlapping regions wi…

Gaussian Processes

Entropy-Tree: Tree-Based Decoding with Entropy-Guided Exploration

2026-01-02 · Longxuan Wei, Yubo Zhang, Zijiao Zhang, Zhihu Wang 외 arxiv

Large language models achieve strong reasoning performance, yet existing decoding strategies either explore blindly (random sampling) or redundantly (independent multi-sampling). We propose Entropy-Tree, a tree-based dec…

On Feature Relevance Uncertainty: A Monte Carlo Dropout Sampling Approach

2020-08-04 · Kai Fischer, Jonas Schneider

Understanding decisions made by neural networks is key for the deployment of intelligent systems in real world applications. However, the opaque decision making process of these systems is a disadvantage where interpreta…

Decision Making