paper-with-me

홈 › Papers

AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching

2025-10-09 · Jingyu Peng, Maolin Wang, Hengyi Cai, Yuchen Li, Kai Zhang, Shuaiqiang Wang, Dawei Yin, Xiangyu Zhao arxiv

Small language models (SLMs) are crucial for applications with strict latency and computational constraints, yet achieving high performance remains challenging. Knowledge distillation (KD) can transfer capabilities from large teacher models, but existing methods face a dilemma: off-policy distillation provides high-quality supervision but suffers from exposure bias (training inference mismatch), while on-policy approaches ensure consistency but are limited by the low quality of student-generated outputs. To address these issues, we propose AdaSwitch, a novel approach that dynamically combines on-policy and off-policy generation via an adaptive switching mechanism. AdaSwitch allows the student to explore its predictions within its capability and selectively integrates teacher guidance only when divergence exceeds a context-aware threshold. This paradigm preserves generation consistency while ensuring high-quality supervision. Experiments on three datasets demonstrate that AdaSwitch consistently improves accuracy and reasoning capability with moderate overhead.

📄 PDF Abstract BibTeX arXiv:2510.07842

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration

2025-10-02 · Xiaoyang Yuan, Yujuan Ding, Yi Bin, Wenqi Shao 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) is a promising paradigm for enhancing the reasoning ability in Large Language Models (LLMs). However, prevailing methods primarily rely on self-exploration or a singl…

Reinforcement LearningKnowledge DistillationMathematical Reasoning

Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models

2026-03-22 · Jingchen Sun, Shaobo Han, Deep Patel, Wataru Kohno 외 arxiv

Knowledge distillation establishes a learning paradigm that leverages both data supervision and teacher guidance. However, determining the optimal balance between learning from data and learning from the teacher is chall…

Knowledge Distillation

Reinforcement-aware Knowledge Distillation for LLM Reasoning

2026-02-26 · Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang 외 arxiv

Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such models motivates distillation into smaller st…

Knowledge DistillationReinforcement Learning

LLaPipe: LLM-Guided Reinforcement Learning for Automated Data Preparation Pipeline Construction

2025-07-18 · Jing Chang, Chang Liu, Jinbin Huang, Rui Mao 외 arxiv

Automated data preparation is crucial for democratizing machine learning, yet existing reinforcement learning (RL) based approaches suffer from inefficient exploration in the vast space of possible preprocessing pipeline…

Computational EfficiencyReinforcement Learning

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

2026-05-05 · Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan 외 arxiv

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model. Despite its empirical success, the con…