paper-with-me

홈 › Papers

Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency

2025-08-06 · Md Arafat Sultan, Ramón Fernandez Astudillo arxiv

Despite its simplicity and efficacy, the high token expenditure of self-consistency can limit its practical utility. Here we investigate if self-consistency can be made more token-efficient for long chain-of-thought reasoning tasks, while preserving its parallelism, through early hypothesis pruning. Concretely, we generate all solutions in parallel, but periodically prune intermediate hypotheses that are deemed unnecessary based on two lightweight indicators: (a) the model's own confidence in individual hypotheses, and (b) lexical coverage of all current hypotheses by candidate subsets that are under consideration for continued retention. We design a fast weighted set cover algorithm that utilizes the two indicators; our evaluation of five LLMs on three math benchmarks shows that this method can improve token efficiency for all models, by 10-35% in many cases.

📄 PDF Abstract BibTeX arXiv:2508.03979

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning

2026-06-01 · Ting Xu, Xu He, Yupu Lu, Jiankai Sun 외 arxiv

This paper investigates the entropy dynamics of Chain-of-Thought (CoT) and uncovers a consistent two-phase structure: an Uncertainty Region of exploration transitioning sharply to a Confidence Region of convergence. We d…

LEAP: Unlocking dLLM Parallelism via Lookahead Early-Convergence Token Detection

2026-05-09 · Haohui Zhang, Zhiye Wang, Xiaoying Gan, Xinbing Wang 외 arxiv

Diffusion Language Models (dLLMs) have garnered significant attention for their potential in highly parallel processing. The parallel capabilities of existing dLLMs stem from the assumption of conditional independence at…

Exact Convex Confidence-Weighted Learning

2008-12-01 · NeurIPS 2008 12 · Koby Crammer, Mark Dredze, Fernando Pereira

Confidence-weighted (CW) learning [6], an online learning method for linear classifiers, maintains a Gaussian distributions over weight vectors, with a covariance matrix that represents uncertainty about weights and corr…

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

2026-08-17 · Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu 외 arxiv

Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managin…

Early-Token Confidence Predicts Reasoning Quality in Multi-Agent LLM Debate

2026-06-09 · Ali Keramati, Justin Cheok, Jacob Horne, Mark Warschauer arxiv

Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whether intrinsic confidence signals, token-level log-probabilities from d…