paper-with-me

홈 › Papers

CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute

2026-02-09 · Chen Jin, Ryutaro Tanno, Tom Diethe, Philip Teare arxiv

Large Language Models (LLMs) often rely on test-time scaling via parallel decoding (for example, 512 samples) to boost reasoning accuracy, but this incurs substantial compute. We introduce CoRefine, a confidence-guided self-refinement method that achieves competitive accuracy using a fraction of the tokens via a lightweight 211k-parameter Conv1D controller atop a frozen LLM. The controller consumes full-trace confidence to decide whether to halt, re-examine, or try a different approach, enabling targeted self-correction with an average of 2.7 refinement steps per problem and roughly 190-fold token reduction relative to 512-sample baselines. Across diverse reasoning benchmarks and three open-source models, the controller achieves 92.6 percent precision when it confidently halts, indicating that confidence dynamics reliably signal correctness without ground-truth verification. We extend this to CoRefine-Tree, a hybrid sequential-parallel variant that adaptively balances exploration and exploitation, with easy serving integration and verifier compatibility. By treating confidence as a control signal rather than a correctness guarantee, CoRefine provides a modular primitive for scalable reasoning and agentic settings with imperfect verifiers.

📄 PDF Abstract BibTeX arXiv:2602.08948

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

2026-07-30 · Hongyu Chen, Liang Lin, Guangrun Wang arxiv

Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refineme…

Mathematical ReasoningReinforcement Learning

Confidence-guided Refinement Reasoning for Zero-shot Question Answering

2025-09-25 · Youwon Jang, Woo Suk Choi, Minjoon Jung, Minsu Lee 외 arxiv

We propose Confidence-guided Refinement Reasoning (C2R), a novel training-free framework applicable to question-answering (QA) tasks across text, image, and video domains. C2R strategically constructs and refines sub-que…

Question Answering

GFSR: Geometric Fidelity and Spatial Refinement for Reliable Lane Detection

2026-05-22 · Tiancheng Wang, Zhaolu Ding, Richeng Xu, Tianhui Zheng 외 arxiv

Lane detection stands as a crucial perception task in autonomous driving and advanced driver assistance systems. However, existing methods still degrade in complex real scenarios due to two major limitations. First, clas…

Autonomous DrivingLane Detection

Discovering New Intents via Constrained Deep Adaptive Clustering with Cluster Refinement

2019-11-20 · Ting-En Lin, Hua Xu, Hanlei Zhang

Identifying new user intents is an essential task in the dialogue system. However, it is hard to get satisfying clustering results since the definition of intents is strongly guided by prior knowledge. Existing methods i…

ClusteringOpen Intent DiscoveryShort Text ClusteringText Clustering

CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency

2025-11-04 · Ehsan Aghazadeh, Ahmad Ghasemi, Hedyeh Beyhaghi, Hossein Pishro-Nik arxiv

Large language models (LLMs) are often queried multiple times at test time, with predictions aggregated by majority vote. While effective, this self-consistency (Wang et al., 2023) strategy requires a fixed number of cal…