paper-with-me

홈 › Papers

Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation

2025-12-24 · Wei-Rui Chen, Vignesh Kothapalli, Ata Fatahibaarzi, Hejian Sang, Shao Tang, Qingquan Song, Zhipeng Wang, Muhammad Abdul-Mageed arxiv

Distilling the capabilities from a large reasoning model (LRM) to a smaller student model often involves training on substantial amounts of reasoning data. However, knowledge distillation (KD) over lengthy sequences with prompt (P), chain-of-thought (CoT), and answer (A) sections makes the process computationally expensive. In this work, we investigate how the allocation of supervision across different sections (P, CoT, A) affects student performance. Our analysis shows that selective KD over only the CoT tokens can be effective when the prompt and answer information is encompassed by it. Building on this insight, we establish a truncation protocol to quantify computation-quality tradeoffs as a function of sequence length. We observe that beyond a specific length, longer training sequences provide marginal returns for downstream performance but require substantially higher memory and FLOPs. To this end, training on only the first $50\%$ of tokens of every training sequence can retain, on average, $\approx91\%$ of full-sequence performance on math benchmarks while reducing training time, memory usage, and FLOPs by about $50\%$ each. Codes are available at https://github.com/weiruichen01/distilling-the-essence.

📄 PDF Abstract BibTeX arXiv:2512.21002

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

Distilling the Implicit Multi-Branch Structure in LLMs' Reasoning via Reinforcement Learning

2025-05-22 · Shicheng Xu, Liang Pang, Yunchang Zhu, Jia Gu 외

Distilling reasoning paths from teacher to student models via supervised fine-tuning (SFT) provides a shortcut for improving the reasoning ability of smaller Large Language Models (LLMs). However, the reasoning paths gen…

Reinforcement Learning (RL)

Reasoning Scaffolding: Distilling the Flow of Thought from LLMs

2025-09-28 · Xiangyu Wen, Junhua Huang, Zeju Li, Min Li 외 arxiv

The prevailing approach to distilling reasoning from Large Language Models (LLMs)-behavioral cloning from textual rationales-is fundamentally limited. It teaches Small Language Models (SLMs) to mimic surface-level patter…

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

2023-06-23 · Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk 외

Knowledge distillation (KD) is widely used for compressing a teacher model to reduce its inference cost and memory footprint, by training a smaller student model. However, current KD methods for auto-regressive sequence …

Arithmetic ReasoningKnowledge DistillationMachine Translation

Hands-on Guidance for Distilling Object Detectors

2021-03-26 · Yangyang Qin, Hefei Ling, Zhenghai He, Yuxuan Shi 외

Knowledge distillation can lead to deploy-friendly networks against the plagued computational complexity problem, but previous methods neglect the feature hierarchy in detectors. Motivated by this, we propose a general f…

Knowledge DistillationObject

Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning

2026-01-14 · Shaotian Yan, Kaiyuan Liu, Chen Shen, Bing Wang 외 arxiv

In this report, we introduce DASD-4B-Thinking, a lightweight yet highly capable, fully open-source reasoning model. It achieves SOTA performance among open-source models of comparable scale across challenging benchmarks …

Code Generation