paper-with-me

Papers

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

2026-08-20 · Gijs Kassenaar, Zhao Yang, Vincent François-Lavet arxiv

Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on difficult ones. We study whether a model can learn to allocate its own reasoning effort by choosing, as the first token of its response, one of three modes: \textsc{NoThink} (answer as quickly as possible), \textsc{Short} (brief reasoning), or \textsc{Long} (extended reasoning). The choice is learned inside Group Relative Policy Optimization (GRPO) with no separate router, through a shaped reward that makes each mode worthwhile at a different response length, together with hard per-mode token caps that keep the modes distinct. On a 1.5B distilled model trained on MATH, the three modes emerge without collapsing to a single choice, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random. Averaged over three seeds, the resulting policy stays close to the base model's accuracy on the held-out MATH500 ($0.782$ vs.\ $0.796$) while cutting the mean response length from $4{,}796$ to $2{,}811$ tokens (a $41\%$ reduction). Interestingly, it also transfers to other benchmarks without retraining, with the largest savings where problems are easier, with for instance 76\% token reduction on GSM8K and at higher accuracy than the baselines at similar response length. In short, we build a reasoning model that adaptively chooses how much to reason for each problem.

📄 PDF Abstract BibTeX arXiv:2608.20256

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Prejudge-Before-Think: Enhancing Large Language Models at Test-Time by Process Prejudge Reasoning

2025-04-18 · Jianing Wang, Jin Jiang, Yang Liu, Mengdi Zhang 외

In this paper, we introduce a new \emph{process prejudge} strategy in LLM reasoning to demonstrate that bootstrapping with process prejudge allows the LLM to adaptively anticipate the errors encountered when advancing th…

Reinforcement Learning (RL)

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks

2026-06-06 · Abid Aziz, Hafsa Binte Kibria arxiv

Reasoning-capable large language models can be induced to spend their generation budget on injected decoy tasks rather than answering the user's question, causing denial of service when no final answer is produced and de…

Question AnsweringCode Generation

Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning

2025-10-11 · Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen 외 arxiv

Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinki…

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

2025-11-03 · Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille 외 arxiv

We propose Re-FORC, an adaptive reward prediction method that, given a context, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapte…

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs

2025-07-02 · Mohammad Ali Alomrani, Yingxue Zhang, Derek Li, Qianyi Sun 외 arxiv

Large language models (LLMs) have rapidly progressed into general-purpose agents capable of solving a broad spectrum of tasks. However, current models remain inefficient at reasoning: they apply fixed inference-time comp…

Computational Efficiency