paper-with-me

홈 › Papers

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

2026-06-29 · Zhe Dong, Fang Qin, Manish Shah arxiv

Reasoning models spend test-time compute unevenly across instances, and a growing family of early-exit rules -- confidence thresholds, entropy monitors, answer-stability checks, and learned stoppers -- promises to reclaim the waste. These rules, however, are evaluated under heterogeneous protocols that leave the deployment question unanswered: at a fixed tolerance for losing correct answers, which policy saves more compute, and does the saving survive probe overhead? We answer this question with a controlled study across 18 task-model settings spanning GSM8K, MATH-500, MMLU-Pro, AIME-90, and GPQA on Qwen3 and DeepSeek-R1-distilled models, using LearnStop, a hidden-state-free logistic stopper over prefix-observable features, as the learned policy instrument. Under matched lost-correct risk at $α$ = 0.15, with the scalar competitor selected on calibration data from confidence, entropy, confidence-leap, and run-stability exits, the answer forms three regimes. Learned stopping wins on all four primary Qwen3 free-form math settings (+3.2 to +21.2 pp additional total-token saving); calibrated scalar exits win on multiple-choice MMLU-Pro; and small hard benchmarks (AIME-90, GPQA) admit no certifiable aggressive policy at all. A trajectory decomposition predicts the regime: learning pays where answers oscillate and correctness evidence is spread across complementary signals, while a single confidence threshold suffices where most instances are already solved at the first checkpoint. Cost accounting sharpens the picture further -- the same policy that saves 32% of tokens under KV-cache forking costs 121% extra under black-box repeated prefilling. Together, these results replace the single-method race with a decision procedure for choosing a stopping rule from the trajectory structure and serving regime of the target workload.

📄 PDF Abstract BibTeX arXiv:2606.30852

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cost-aware Stopping for Bayesian Optimization

2025-07-16 · Qian Xie, Linda Cai, Alexander Terenin, Peter I. Frazier 외 arxiv

In automated machine learning, scientific discovery, and other applications of Bayesian optimization, deciding when to stop evaluating expensive black-box functions in a cost-aware manner is an important but underexplore…

Hyperparameter Optimization

ACE: Adaptive Constraint-aware Early Stopping in Hyperparameter Optimization

2022-08-04 · Yi-Wei Chen, Chi Wang, Amin Saied, Rui Zhuang

Deploying machine learning models requires high model quality and needs to comply with application constraints. That motivates hyperparameter optimization (HPO) to tune model configurations under deployment constraints. …

FairnessHyperparameter Optimization

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

2026-07-29 · Yicheng Feng, Yan Zhang, Yan Cheng, Wei Qi arxiv

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under…

AdaStop: Cost-Aware Early Stopping for DNN Test Selection

2026-07-06 · Bonan Shen, Wei-Jung Huang, Xin Liu, Jiazhou Gao 외 arxiv

Existing methods for testing deep neural networks (DNNs) primarily prioritize test inputs likely to reveal model faults under a fixed labeling budget. In practice, choosing that budget is difficult: too little testing mi…

Competition, Persuasion, and Search

2024-11-17 · Teddy Mekonnen, Bobak Pakzad-Hurson

An agent engages in sequential search. He does not directly observe the quality of the goods he samples, but he can purchase signals designed by profit maximizing principal(s). We formulate the principal-agent relationsh…