paper-with-me

홈 › Papers

Variational Information Distillation for Knowledge Transfer

2019-04-11 · CVPR 2019 6 · Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D. Lawrence, Zhenwen Dai

Transferring knowledge from a teacher neural network pretrained on the same or a similar task to a student neural network can significantly improve the performance of the student neural network. Existing knowledge transfer approaches match the activations or the corresponding hand-crafted features of the teacher and the student networks. We propose an information-theoretic framework for knowledge transfer which formulates knowledge transfer as maximizing the mutual information between the teacher and the student networks. We compare our method with existing knowledge transfer methods on both knowledge distillation and transfer learning tasks and show that our method consistently outperforms existing methods. We further demonstrate the strength of our method on knowledge transfer across heterogeneous network architectures by transferring knowledge from a convolutional neural network (CNN) to a multi-layer perceptron (MLP) on CIFAR-10. The resulting MLP significantly outperforms the-state-of-the-art methods and it achieves similar performance to the CNN with a single convolutional layer.

📄 PDF Abstract BibTeX arXiv:1904.05835

Code (2)

amzn/xfer mxnet
yoshitomo-matsubara/torchdistill pytorch

Tasks

Knowledge DistillationTransfer Learning

Similar Papers 제목 키워드 기반

Entire-Space Variational Information Exploitation for Post-Click Conversion Rate Prediction

2024-12-17 · Ke Fei, Xinyue Zhang, Jingjing Li

In recommender systems, post-click conversion rate (CVR) estimation is an essential task to model user preferences for items and estimate the value of recommendations. Sample selection bias (SSB) and data sparsity (DS) a…

Knowledge DistillationRecommendation SystemsSelection bias

Learning to Maximize Mutual Information for Chain-of-Thought Distillation

2024-03-05 · Xin Chen, Hanxian Huang, Yanjun Gao, Yi Wang 외

Knowledge distillation, the technique of transferring knowledge from large, complex models to smaller ones, marks a pivotal step towards efficient AI deployment. Distilling Step-by-Step~(DSS), a novel method utilizing ch…

Knowledge DistillationLanguage ModelingLanguage ModellingMulti-Task Learning

A Unified Knowledge Distillation Framework for Deep Directed Graphical Models

2021-09-29 · CVPR 2023 1 · Yizhuo Chen, Kaizhao Liang, Zhe Zeng, Yifei Yang 외

Knowledge distillation (KD) is a technique that transfers the knowledge from a large teacher network to a small student network. It has been widely applied to many different tasks, such as model compression and federate…

Continual LearningFederated LearningKnowledge DistillationModel Compression

Variational Knowledge Distillation for Disease Classification in Chest X-Rays

2021-03-19 · Tom van Sonsbeek, XianTong Zhen, Marcel Worring, Ling Shao

Disease classification relying solely on imaging data attracts great interest in medical image analysis. Current models could be further improved, however, by also employing Electronic Health Records (EHRs), which contai…

ClassificationGeneral Classificationimage-classificationImage Classification+3

Swing Distillation: A Privacy-Preserving Knowledge Distillation Framework

2022-12-16 · Junzhuo Li, Xinwei Wu, Weilong Dong, Shuangzhi Wu 외

Knowledge distillation (KD) has been widely used for model compression and knowledge transfer. Typically, a big teacher model trained on sufficient data transfers knowledge to a small student model. However, despite the …

Knowledge DistillationModel CompressionPrivacy PreservingTransfer Learning