paper-with-me

홈 › Papers

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

2025-08-13 · Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai, Dengfeng Liu, Qiaobo Hao, Huang Liu, Zhiguo Yang, Ji Xie, Ninglun Gu, Jin Yang, Kailai Zhang, Yelun Bao, Jun Wang arxiv

Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving. Recent methods have improved reasoning through expanded corpus and multistage training combining reinforcement learning and supervised fine-tuning. Although some methods suggest that small but targeted dataset can incentivize reasoning via only distillation, a reasoning scaling laws is still taking shape, increasing computational costs. To address this, we propose a data-efficient distillation framework (DED) that optimizes the Pareto frontier of reasoning distillation. Inspired by the on-policy learning and diverse roll-out strategies of reinforcement learning, the key idea of our approach is threefold: (1) We identify that benchmark scores alone do not determine an effective teacher model. Through comprehensive comparisons of leading reasoning LLMs, we develop a method to select an optimal teacher model. (2) While scaling distillation can enhance reasoning, it often degrades out-of-domain performance. A carefully curated, smaller corpus achieves a balanced trade-off between in-domain and out-of-domain capabilities. (3) Diverse reasoning trajectories encourage the student model to develop robust reasoning skills. We validate our method through evaluations on mathematical reasoning (AIME 2024/2025, MATH-500) and code generation (LiveCodeBench), achieving state-of-the-art results with only 0.8k carefully curated examples, bypassing the need for extensive scaling. Our systematic analysis demonstrates that DED outperforms existing methods by considering factors beyond superficial hardness, token length, or teacher model capability. This work offers a practical and efficient pathway to advanced reasoning while preserving general capabilities.

📄 PDF Abstract BibTeX arXiv:2508.09883

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical ReasoningCode Generation

Similar Papers 제목 키워드 기반

Reasoning Beyond Limits: Advances and Open Problems for LLMs

2025-03-26 · Mohamed Amine Ferrag, Norbert Tihanyi, Merouane Debbah

Recent generative reasoning breakthroughs have transformed how large language models (LLMs) tackle complex problems by dynamically retrieving and refining information while generating coherent, multi-step thought process…

Mixture-of-ExpertsRAGreinforcement-learningReinforcement Learning+3

Scaling Reasoning Efficiently via Relaxed On-Policy Distillation

2026-03-11 · Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen 외 arxiv

On-policy distillation is pivotal for transferring reasoning capabilities to capacity-constrained models, yet remains prone to instability and negative transfer. We show that on-policy distillation can be interpreted, bo…

Visual Reasoning

NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions

2025-02-18 · Weizhe Yuan, Jane Yu, Song Jiang, Karthik Padthe 외

Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generatin…

Knowledge DistillationMath

The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models

2025-10-07 · Muyu He, Muhammad Ali Shafique, Anand Kumar, Tsach Mackey 외 arxiv

Distilling the thinking traces of a Large Language Model (LLM) with reasoning capabilities into a smaller model has been proven effective. Yet, there is a scarcity of work done on how model performances scale with the qu…

Knowledge Distillation

Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

2026-02-12 · Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang 외 arxiv

On-policy distillation (OPD), which aligns the student with the teacher's logit distribution on student-generated trajectories, has demonstrated strong empirical gains in improving student performance and often outperfor…

Reinforcement LearningCode Generation