paper-with-me

홈 › Papers

When RL Suppresses Its Own Vocabulary: Recovering Reasoning Diversity in Puzzle-to-Math Transfer

2026-05-28 · Mayug Maniparambil, Arjun Karuvally, Terrence Sejnowski, Fergal Reid arxiv

Reinforcement learning using verifiable rewards (RLVR) improves LLM reasoning, but the conditions under which it transfers across domains -- and why it does so -- remain under-explored. We study cross-domain transfer in a 7B model whose SFT and RL post-training stages use only constraint-satisfaction puzzles, with no mathematics problems in the post-training data. To analyze how transfer emerges, we introduce a reasoning primitive-level framework that combines a 9-class span classifier with motif extraction, allowing us to segment chain-of-thought traces into primitive motifs and track their evolution across training stages and domains. We find that puzzle SFT induces a reasoning-primitive vocabulary, yielding a $+7$pp \texttt{pass@32} gain on OlymMATH-Hard. Vanilla GSPO then composes these primitives into longer compute-verify chains, adding a further $+6$pp. However, this RL stage also suppresses exploratory primitives such as \textit{hypothesize} and \textit{backtrack}. To address this, we introduce a novelty bonus that rewards diverse correct rollouts, using perplexity under the reference model as a signal. This restores recovery primitives during RL and adds a further $+7$pp \texttt{pass@32} relative to vanilla GSPO. Finally, the end-to-end recipe raises the hard-math capability ceiling from $16.0\%$ at the OLMo3-7B-Instruct-SFT base to $36.0\%$, without adding any mathematics problems during the SFT or RL stages.

📄 PDF Abstract BibTeX arXiv:2605.29190

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models

2025-10-11 · Jinbin Zhang, Nasib Ullah, Erik Schultheis, Rohit Babbar arxiv

Speculative decoding accelerates LLM inference by letting a small drafter propose multiple tokens which a large target model verifies once per speculation step. As vocabularies scale past 10e5 tokens,verification cost in…

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

2026-06-21 · Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher 외 arxiv

In open-ended generation, LLMs frequently fall into the "likelihood trap", marked by repetitive degeneration and vocabulary dullness, creating a discrepancy between machine-generated and human-written text. While post-ho…

Mathematical Reasoning

Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?

2026-02-07 · Alexander von Recum, Leander Girrbach, Zeynep Akata arxiv

Reasoning LLMs (RLLMs) generate step-by-step chains of thought (CoTs) before giving an answer, which improves performance on complex tasks and makes reasoning more transparent. But how robust are these reasoning traces t…

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

2026-05-11 · Jeonghye Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang arxiv

Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the…

Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models

2025-11-06 · Saurabh Srivastava, Janit Bidhan, Hao Yan, Abhishek Dey 외 arxiv

Large Reasoning Models (LRMs) achieve strong performance through explicit chain-of-thought reasoning but suffer from \textit{overthinking}: generating excessive reasoning tokens even for trivial queries. {Beyond inflatin…