paper-with-me

홈 › Papers

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation

2025-09-26 · Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad, Céline Hudelot, Pierre Colombo arxiv

Distilling reasoning traces from strong teacher models has become the standard recipe for building capable small language models. Yet reasoning traces are 5-20$\times$ longer than standard instruction fine-tuning (IFT) outputs, meaning every practitioner who chooses reasoning distillation implicitly forgoes training a larger IFT model on the same compute budget. Whether this trade-off is worthwhile remains unaddressed. We study it with a controlled experiment: a single teacher generates paired IFT and reasoning outputs for identical prompts by toggling only its reasoning mode, isolating supervision format as the sole variable. Training students at five scales (0.5B to 14B) and evaluating on 18 benchmarks, we find that at matched FLOPs, IFT lies on or near the Pareto frontier across the majority of configurations. Reasoning reaches the Pareto frontier only on open-ended tasks at 7B and above. Even there, a sequential curriculum mixing just 25-50\% reasoning data with IFT captures most of the accuracy benefit at far lower compute cost.

📄 PDF Abstract BibTeX arXiv:2509.22193

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

2025-02-07 · Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 외

We study a novel language model architecture that is capable of scaling test-time computation by implicitly reasoning in latent space. Our model works by iterating a recurrent block, thereby unrolling to arbitrary depth …

Language ModelingLanguage Modelling

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

2026-04-29 · Zhenyu Zhao, Sander Land, Daniel M. Bikel, Waseem Alshikh arxiv

Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains underexplored. We observe that reasoning tokens split into two functiona…

Mathematical Reasoning

The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

2026-08-18 · Eduardo Sánchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee 외 arxiv

Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We prese…

Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning

2026-05-20 · Benhao Huang, Zhengyang Geng, Zico Kolter arxiv

Scaling test-time compute by iteratively updating a latent state has emerged as a powerful paradigm for reasoning. Yet the internal mechanisms that enable these iterative models to generalize beyond memorized patterns re…

ThinkBrake: Efficient Reasoning via Log-Probability Margin Guided Decoding

2025-10-01 · Sangjun Song, Minjae Oh, Seungkyu Lee, Sungmin Jo 외 arxiv

Large Reasoning Models (LRMs) allocate substantial inference-time compute to Chain-of-Thought (CoT) reasoning, improving performance on mathematics, scientific QA, and tool usage. However, this introduces overthinking: L…