paper-with-me

홈 › Papers

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

2026-07-30 · Hongyu Chen, Liang Lin, Guangrun Wang arxiv

Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues refinement using its own self-verification. Ground-truth correctness is used only to construct training rewards and is never exposed to the policy through refinement prompts or required at inference. SVR is trained with GRPO on fixed-horizon trajectories using rewards that promote solution correctness, calibration-aware self-verification, and stop-ready correct states; adaptive stopping is activated only at inference. On seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves a macro-average accuracy of 0.563 with only 2.99 inference turns on average. In the evaluated complete-system comparison, it exceeds standard GRPO, strong multi-turn baselines, and a fixed-budget oracle-guided score-feedback reference while requiring substantially fewer turns than fixed ten-turn inference. These results demonstrate that learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute allocation.

📄 PDF Abstract BibTeX arXiv:2607.28457

Code (2)

Aaron617/agent-arXiv-daily ★ 10
Tavish9/awesome-daily-AI-arxiv ★ 112

Tasks

Mathematical ReasoningReinforcement Learning

Similar Papers 제목 키워드 기반

TSVer: A Benchmark for Fact Verification Against Time-Series Evidence

2025-11-02 · Marek Strong, Andreas Vlachos arxiv

Reasoning over temporal and numerical data, such as time series, is a crucial aspect of fact-checking. While many systems have recently been developed to handle this form of evidence, their evaluation remains limited by …

Fact Verification

CAMEL: Confidence-Gated Reflection for Reward Modeling

2026-02-24 · Zirui Zhu, Hailun Xu, Yang Luo, Yong Liu 외 arxiv

Reward models play a fundamental role in aligning large language models with human preferences. Existing methods predominantly follow two paradigms: scalar discriminative preference models, which are efficient but lack i…

Reinforcement Learning

Toward `verifying' a Water Treatment System

2017-12-12 · Jingyi Wang, Jun Sun, Yifan Jia, Shengchao Qin 외

Modeling and verifying real-world cyber-physical systems is challenging, which is especially so for complex systems where manually modeling is infeasible. In this work, we report our experience on combining model learnin…

ssVERDICT: Self-Supervised VERDICT-MRI for Enhanced Prostate Tumour Characterisation

2023-09-12 · Snigdha Sen, Saurabh Singh, Hayley Pye, Caroline M. Moore 외

Purpose: Demonstrating and assessing self-supervised machine learning fitting of the VERDICT (Vascular, Extracellular and Restricted DIffusion for Cytometry in Tumours) model for prostate. Methods: We derive a self-super…

Diffusion MRISelf-Supervised Learning

AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web

2023-05-22 · NeurIPS 2023 11 · Michael Schlichtkrull, Zhijiang Guo, Andreas Vlachos

Existing datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the cla…

Claim VerificationFact CheckingQuestion Answering