paper-with-me

홈 › Papers

CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency

2025-11-04 · Ehsan Aghazadeh, Ahmad Ghasemi, Hedyeh Beyhaghi, Hossein Pishro-Nik arxiv

Large language models (LLMs) are often queried multiple times at test time, with predictions aggregated by majority vote. While effective, this self-consistency (Wang et al., 2023) strategy requires a fixed number of calls and fails when the correct answer is infrequent. We introduce Confidence-Guided Early Stopping (CGES), a Bayesian framework that forms posteriors over candidate answers and adaptively halts sampling once one answer accumulates enough posterior mass. We prove guarantees in both an ideal calibrated regime and a realistic noisy-confidence regime under a directional drift condition. Averaged over five reasoning benchmarks, CGES reduces the average number of calls by 58% on average (from 16.0 to 6.7) while matching its accuracy within 0.4 percentage points of self-consistency.

📄 PDF Abstract BibTeX arXiv:2511.02603

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Second-order Confidence Network for Early Classification of Time Series

2023-12-19 · ACM Transactions on Intelligent Systems and Technology 2023 12 · Junwei Lv, Yuqi Chu, Jun Hu, Peipei Li 외

Time series data are ubiquitous in a variety of disciplines. Early classification of time series, which aims to predict the class label of a time series as early and accurately as possible, is a significant but challengi…

Early ClassificationTime Series

Early Stopping for Large Reasoning Models via Confidence Dynamics

2026-04-06 · Parsa Hosseini, Sumit Nawathe, Mahdi Salmani, Meisam Razaviyayn 외 arxiv

Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key ch…

Reflective Confidence: Correcting Reasoning Flaws via Online Self-Correction

2025-12-21 · Qinglin Zeng, Jing Yang, Keze Wang arxiv

Large language models (LLMs) have achieved strong performance on complex reasoning tasks using techniques such as chain-of-thought and self-consistency. However, ensemble-based approaches, especially self-consistency whi…

Mathematical Reasoning

ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning

2025-05-08 · Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang 외

Large Reasoning Models (LRMs) perform strongly in complex reasoning tasks via Chain-of-Thought (CoT) prompting, but often suffer from verbose outputs caused by redundant content, increasing computational overhead, and de…

Early Stopping Based on Repeated Significance

2024-08-01 · Eric Bax, Arundhyoti Sarkar, Alex Shtoff

For a bucket test with a single criterion for success and a fixed number of samples or testing period, requiring a $p$-value less than a specified value of $\alpha$ for the success criterion produces statistical confiden…