paper-with-me

홈 › Papers

How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum

2026-04-28 · Chu-Cheng Lin, Eugene Ie arxiv

SFT-then-RLVR is widely used for post-training reasoning models, but why this specific ordering, and why RLVR-only stalls at cold start, have lacked a unifying theoretical account. We provide that account under a unified loss family $J_Q$ using the Tsallis $q$-logarithm. $J_Q$ is a single-parameter family that interpolates between RLVR (at $q{=}0$, the \textit{exploitation pole}) and the log-marginal-likelihood over latent trajectories (at $q{=}1$, the \textit{density-estimation pole}), under which the standard pipeline corresponds to a stepwise $q{=}1 \to 0$ schedule. All members share the same per-example gradient direction, differing only by a per-instance amplification $P_θ^{-q}$ that reweights each instance independently of the learning rate. Under gradient flow analysis, we show that the exploitation pole requires $Ω(\frac{1}{p_0})$ time to escape cold start but is robust to label noise, while the density-estimation pole escapes in $Θ\big(\log(\frac{1}{p_0})\big)$ but memorizes label noise. This separation explains how SFT ($q{=}1$) first moves the model out of the cold-start regime, followed by the more robust RLVR ($q{=}0$), under the SFT-then-RLVR paradigm. We further derive two Monte Carlo estimators that directly optimize fixed-$q$ on the $J_Q$ continuum, without annotated rationales: Gradient-Amplified RL (GARL) and Posterior-Attenuated Fine-Tuning (PAFT), with shared bias $O\big(\frac{q}{M P_θ^q}\big)$ but different variance and stability properties. On FinQA, HotPotQA, and MuSiQue, GARL at sufficiently high $q$ substantially mitigates cold-start stalling, escaping cold start where GRPO fails entirely. In warm start, GARL at low $q$ dominates FinQA where training is stable; on HotPotQA and MuSiQue, GARL destabilizes and PAFT at $q{=}0.75$ remains stable, reaching $47.9$ \texttt{m@16} on HotPotQA ($+13.9$ over GRPO).

📄 PDF Abstract BibTeX arXiv:2604.25907

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Path Matters: Learning a Token-Commitment Policy for Diffusion Language Models

2026-05-23 · Bohang Sun, Max Zhu, Francesco Caso, Jindong Gu 외 arxiv

Diffusion large language models promise faster generation by refining many token positions in parallel, but this parallelism introduces a hidden control problem: which proposed tokens should be transferred into the parti…

Mathematical ReasoningQuestion AnsweringCode Generation

State commitment learning: training language models to distinguish computation from memory

2026-05-22 · Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao 외 arxiv

Reasoning language models do not distinguish tokens used for computation from tokens that constitute persistent state: once generated, all hidden thoughts remain in context and influence future predictions. As a result, …

Pass the Baton: Trajectory-Relayed On-Policy Distillation

2026-07-28 · Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni 외 arxiv

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on …

Mathematical Reasoning

Capacity-Dependent Effects of Data Selection for Reasoning

2026-08-13 · Cuong Dang, Hoang Anh Just, Ruoxi Jia arxiv

In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods sug…

Mathematical Reasoning

Learning under noisy supervision is governed by a feedback-truth gap

2026-02-18 · Elan Schonfeld, Elias Wisnia arxiv

When feedback is absorbed faster than task structure can be evaluated, the learner will favor feedback over truth. A two-timescale model shows this feedback-truth gap is inevitable whenever the two rates differ and vanis…