paper-with-me

홈 › Papers

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

2026-08-17 · Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu, Haotian Lin, Yuling Shi, Min Wang, Beijun Shen arxiv

Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these parallel paths typically rely on final-answer consensus, local token confidence, or isolated intermediate probes. However, these signals are often delayed, weakly tied to actual reasoning progress, or too noisy for dynamic, branch-level control. To address these limitations, we introduce ParaTempo, a training-free asynchronous parallel reasoning framework. ParaTempo is driven by temporal confidence, a branch-local measure of answer-space convergence. Each branch is periodically probed for a tentative answer probability distribution, and temporal confidence quantifies how sharply the recent intermediate probes concentrate on a dominant answer. Once sufficient evidence has accumulated, ParaTempo drives its entire control process from this single signal: low-confidence branches are pruned, branches that persistently commit to their dominant answer are retired early, freed computation is reallocated by forking new branches, and generation stops globally once the confidence-weighted vote concentrates. Without requiring synchronization among reasoning trajectories, ParaTempo adaptively allocates computation based on branch-level convergence. Experiments on challenging mathematical and scientific reasoning benchmarks show that ParaTempo reduces average latency by 21.8-32.2% and total token usage by 18.1-30.3% while maintaining competitive accuracy. Moreover, temporal confidence exhibits stronger temporal stability and predictive power for future branch convergence than token-level and instantaneous signals.

📄 PDF Abstract BibTeX arXiv:2608.16425

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning

2026-01-19 · Zhenjiang Mao, Anirudhh Venkat, Artem Bisliouk, Akshat Kothiyal 외 arxiv

Large Language Models (LLMs) increasingly rely on long-form, multi-step reasoning to solve complex tasks such as mathematical problem solving and scientific question answering. Despite strong performance, existing confid…

Question Answering

Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic

2025-06-09 · Zhenjiang Mao, Artem Bisliouk, Rohith Reddy Nama, Ivan Ruchkin

Large Language Models (LLMs) have shown impressive performance in mathematical reasoning tasks when guided by Chain-of-Thought (CoT) prompting. However, they tend to produce highly confident yet incorrect outputs, which …

Mathematical Reasoning

Selective Temporal Knowledge Graph Reasoning

2024-04-02 · Zhongni Hou, Xiaolong Jin, Zixuan Li, Long Bai 외

Temporal Knowledge Graph (TKG), which characterizes temporally evolving facts in the form of (subject, relation, object, timestamp), has attracted much attention recently. TKG reasoning aims to predict future facts based…

CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute

2026-02-09 · Chen Jin, Ryutaro Tanno, Tom Diethe, Philip Teare arxiv

Large Language Models (LLMs) often rely on test-time scaling via parallel decoding (for example, 512 samples) to boost reasoning accuracy, but this incurs substantial compute. We introduce CoRefine, a confidence-guided s…

Efficient Diffusion LLMs via Temporal-Spatial Parallel Decoding and Confidence Extrapolation

2026-05-29 · Zekai Li, Ji Liu, Yiqing Huang, Ziqiong Liu 외 arxiv

Diffusion-based large language models (dLLMs) support parallel text generation via iterative denoising, yet inference remains latency-heavy because many steps are spent on redundant refinement and repeated remasking of t…

Text Generation