paper-with-me

홈 › Papers

SkillFactory: Self-Distillation For Learning Cognitive Behaviors

2025-12-03 · Zayne Sprague, Jack Lu, Manya Wadhwa, Sedrick Keh, Mengye Ren, Greg Durrett arxiv

Reasoning models leveraging long chains of thought employ various cognitive skills, such as verification of their answers, backtracking, retrying by an alternate method, and more. Previous work has shown that when a base language model exhibits these skills, training that model further with reinforcement learning (RL) can learn to leverage them. How can we get models to leverage skills that aren't exhibited by base models? Our work, SkillFactory, is a method for fine-tuning models to roughly learn these skills during a supervised fine-tuning (SFT) stage prior to RL. Our approach does not rely on distillation from a stronger model, but instead uses samples from the model itself, rearranged to provide training data in the format of those skills. These "silver" SFT traces may be imperfect, but are nevertheless effective for priming a model to acquire skills during RL. Our evaluation shows that (1) starting from SkillFactory SFT initialization helps a model to generalize to harder variants of a task post-RL, despite lower performance pre-RL;(2) cognitive skills are indeed used by the model; (3) RLed SkillFactory models are more robust to regression on out-of-domain tasks than RLed base models. Our work suggests that inductive biases learned prior to RL help models learn robust cognitive skill use.

📄 PDF Abstract BibTeX arXiv:2512.04072

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

2026-07-08 · Xinyu Geng, Xuanhua He, Sixiang Chen, Yanjing Xiao 외 arxiv

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak superv…

Reinforcement Learning

Leave No One Behind: Online Self-Supervised Self-Distillation for Sequential Recommendation

2024-03-22 · Shaowei Wei, Zhengwei Wu, Xin Li, Qintong Wu 외

Sequential recommendation methods play a pivotal role in modern recommendation systems. A key challenge lies in accurately modeling user preferences in the face of data sparsity. To tackle this challenge, recent methods …

ClusteringContrastive LearningOnline ClusteringRecommendation Systems+2

ThinkTuning: Instilling Cognitive Reflections without Distillation

2025-08-11 · Aswin RRV, Jacob Dineen, Divij Handa, Md Nayem Uddin 외 arxiv

Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning. While RL drives this self-improvement paradigm, a recent study (Gandhi et a…

Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

2025-03-03 · Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile 외

Test-time inference has emerged as a powerful paradigm for enabling language models to ``think'' longer and more carefully about complex challenges, much like skilled human experts. While reinforcement learning (RL) can …

Reinforcement Learning (RL)

"Favoring my playmate seems fair": Inhibitory control and theory of mind in preschoolers' self-disadvantaging behaviors

2019-04-25

The purpose of this study was to investigate the relationship between preschoolers' cognitive abilities and their fairness-related allocation behaviors in a dilemma of equity-efficiency conflict. Four- to 6-year-olds in …

Fairness