paper-with-me

홈 › Papers

When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling

2026-04-29 · Zhimin Lin, Yixin Ji, Jinpeng Li, Yu Luo, Dong Li, Junhua Fang, Juntao Li, Min Zhang arxiv

Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-time scaling methods, such as repeated sampling, self-correction, and tree search, improve performance at the cost of increased computation, yet often exhibit diminishing returns on hard problems. We observe that output disagreement is strongly correlated with instance difficulty and prediction correctness, providing a useful signal for guiding instance-level strategy selection at test time. Based on this insight, we propose a training-free framework that formulates test-time scaling as an instance-level routing problem, rather than allocating more computation within a single strategy, dynamically selecting among different scaling strategies based on output disagreement. The framework applies lightweight resolution for consistent cases, majority voting for moderate disagreement, and rewriting-based reformulation for highly ambiguous instances. Experiments on seven mathematical benchmarks and three models show that our method improves accuracy by 3% - 7% while reducing sampling cost compared to existing approaches.

📄 PDF Abstract BibTeX arXiv:2604.26644

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Vote-boosting ensembles

2016-06-30 · Maryam Sabzevari, Gonzalo Martínez-Muñoz, Alberto Suárez

Vote-boosting is a sequential ensemble learning method in which the individual classifiers are built on different weighted versions of the training data. To build a new classifier, the weight of each training instance is…

Ensemble Learning

Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome Classification

2024-02-11 · Shanshan Xu, T. Y. S. S Santosh, Oana Ichim, Barbara Plank 외

In legal decisions, split votes (SV) occur when judges cannot reach a unanimous decision, posing a difficulty for lawyers who must navigate diverse legal arguments and opinions. In high-stakes domains, understanding the …

Navigate

Domain Adaptation of Majority Votes via Perturbed Variation-based Label Transfer

2013-11-19 · Emilie Morvant

We tackle the PAC-Bayesian Domain Adaptation (DA) problem. This arrives when one desires to learn, from a source distribution, a good weighted majority vote (over a set of classifiers) on a different target distribution.…

Domain Adaptation

Politicians' Willingness to Agree: Evidence from the interactions in Twitter of Chilean Deputies

2021-06-16 · Pablo Henríquez, Jorge Sabat, José Patrìcio Sullivan

Measuring the number of "likes" in Twitter and the number of bills voted in favor by the members of the Chilean Chambers of Deputies. We empirically study how signals of agreement in Twitter translates into cross-cutting…

Same Target, Different Basins: Hard vs. Soft Labels for Annotator Distributions

2026-05-20 · Mirerfan Gheibi, Gashin Ghazizadeh arxiv

When annotators disagree, that disagreement can reflect epistemic uncertainty rather than simple label noise. We study hard-label delivery as an alternative to the usual choices of collapsing votes to a single label or t…