paper-with-me

홈 › Papers

NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks

2025-07-02 · Yang Li, Youssef Emad, Karthik Padthe, Jack Lanchantin, Weizhe Yuan, Thao Nguyen, Jason Weston, Shang-Wen Li, Dong Wang, Ilia Kulikov, Xian Li arxiv

Recent work has shown that distilling reasoning traces from a larger teacher model via supervised finetuning outperforms reinforcement learning with the smaller student model alone (Guo et al. 2025). However, there has not been a systematic study of what kind of reasoning demonstrations from the teacher are most effective in improving the student model's reasoning capabilities. In this work we curate high-quality "NaturalThoughts" by selecting reasoning traces from a strong teacher model based on a large pool of questions from NaturalReasoning (Yuan et al. 2025). We first conduct a systematic analysis of factors that affect distilling reasoning capabilities, in terms of sample efficiency and scalability for general reasoning tasks. We observe that simply scaling up data size with random sampling is a strong baseline with steady performance gains. Further, we find that selecting difficult examples that require more diverse reasoning strategies is more sample-efficient to transfer the teacher model's reasoning skills. Evaluated on both Llama and Qwen models, training with NaturalThoughts outperforms existing reasoning datasets such as OpenThoughts, LIMO, etc. on general STEM reasoning benchmarks including GPQA-Diamond, MMLU-Pro and SuperGPQA.

📄 PDF Abstract BibTeX arXiv:2507.01921

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

The Signal is in the Steps: Local Scoring for Reasoning Data Selection

2025-10-05 · Hoang Anh Just, Myeongseob Ko, Ruoxi Jia arxiv

Distilling long-form reasoning from teacher models into smaller students requires selecting which candidate solutions to train on. Recent work argues that one should select responses the student model assigns highest pro…

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding

2026-05-04 · Taewon Yun, Jisu Shin, Jeonghwan Choi, Seunghwan Bang 외 arxiv

Distilling large reasoning models is essential for making Long-CoT reasoning practical, as full-scale inference remains computationally prohibitive. Existing curation-based approaches select complete reasoning traces pos…

Efficient Reasoning on the Edge

2026-03-17 · Yelysei Bondarenko, Thomas Hehn, Rob Hesselink, Romain Lepert 외 arxiv

Large language models (LLMs) with chain-of-thought reasoning achieve state-of-the-art performance across complex problem-solving tasks, but their verbose reasoning traces and large context requirements make them impracti…

Reinforcement Learning

Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning

2025-04-06 · Ximing Lu, Seungju Han, David Acuna, Hyunwoo Kim 외

Large reasoning models exhibit remarkable reasoning capabilities via long, elaborate reasoning trajectories. Supervised fine-tuning on such reasoning traces, also known as distillation, can be a cost-effective way to boo…

Math

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation

2025-09-26 · Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad, Céline Hudelot 외 arxiv

Distilling reasoning traces from strong teacher models has become the standard recipe for building capable small language models. Yet reasoning traces are 5-20$\times$ longer than standard instruction fine-tuning (IFT) o…