paper-with-me

홈 › Papers

Hybrid Data-Free Knowledge Distillation

2024-12-18 · Jialiang Tang, Shuo Chen, Chen Gong

Data-free knowledge distillation aims to learn a compact student network from a pre-trained large teacher network without using the original training data of the teacher network. Existing collection-based and generation-based methods train student networks by collecting massive real examples and generating synthetic examples, respectively. However, they inevitably become weak in practical scenarios due to the difficulties in gathering or emulating sufficient real-world data. To solve this problem, we propose a novel method called \textbf{H}ybr\textbf{i}d \textbf{D}ata-\textbf{F}ree \textbf{D}istillation (HiDFD), which leverages only a small amount of collected data as well as generates sufficient examples for training student networks. Our HiDFD comprises two primary modules, \textit{i.e.}, the teacher-guided generation and student distillation. The teacher-guided generation module guides a Generative Adversarial Network (GAN) by the teacher network to produce high-quality synthetic examples from very few real-world collected examples. Specifically, we design a feature integration mechanism to prevent the GAN from overfitting and facilitate the reliable representation learning from the teacher network. Meanwhile, we drive a category frequency smoothing technique via the teacher network to balance the generative training of each category. In the student distillation module, we explore a data inflation strategy to properly utilize a blend of real and synthetic data to train the student network via a classifier-sharing-based feature alignment technique. Intensive experiments across multiple benchmarks demonstrate that our HiDFD can achieve state-of-the-art performance using 120 times less collected data than existing methods. Code is available at https://github.com/tangjialiang97/HiDFD.

📄 PDF Abstract BibTeX arXiv:2412.13525

Code (1)

tangjialiang97/hidfd 공식 구현 pytorch

Tasks

Data-free Knowledge DistillationGenerative Adversarial NetworkKnowledge DistillationRepresentation Learning

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models

2026-03-27 · Juan Gabriel Kostelec, Xiang Wang, Axel Laborieux, Christos Sourmpis 외 arxiv

Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs. However, achieving high-quality generation in distilled models requires…

Reviving Stale Updates: Data-Free Knowledge Distillation for Asynchronous Federated Learning

2025-11-01 · Baris Askin, Holger R. Roth, Zhenyu Sun, Carlee Joe-Wong 외 arxiv

Federated learning (FL) enables collaborative model training across distributed clients without sharing raw data, yet its scalability is limited by synchronization overhead. Asynchronous federated learning (AFL) alleviat…

Data-free Knowledge DistillationFederated Learning

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

2026-08-12 · Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen 외 hf

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates know…

knowledge editing

Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation

2024-04-23 · Jingxuan Wei, Linzhuang Sun, Yichong Leng, Xu Tan 외

Knowledge distillation, transferring knowledge from a teacher model to a student model, has emerged as a powerful technique in neural machine translation for compressing models or simplifying training targets. Knowledge …

Knowledge DistillationMachine TranslationSentence

Staged Hybridisation for Visual Quantum Reinforcement Learning via Knowledge Distillation

2026-06-29 · Javier Lazaro, Juan-Ignacio Vazquez, Pablo Garcia-Bringas arxiv

Visual environments are a demanding setting for quantum reinforcement learning (QRL): high-dimensional observations, unstable RL optimisation, and constrained variational quantum circuits (VQCs) are difficult to train jo…

Reinforcement LearningKnowledge Distillation