paper-with-me

홈 › Papers

Preparing Lessons: Improve Knowledge Distillation with Better Supervision

2019-11-18 · Tiancheng Wen, Shenqi Lai, Xueming Qian

Knowledge distillation (KD) is widely used for training a compact model with the supervision of another large model, which could effectively improve the performance. Previous methods mainly focus on two aspects: 1) training the student to mimic representation space of the teacher; 2) training the model progressively or adding extra module like discriminator. Knowledge from teacher is useful, but it is still not exactly right compared with ground truth. Besides, overly uncertain supervision also influences the result. We introduce two novel approaches, Knowledge Adjustment (KA) and Dynamic Temperature Distillation (DTD), to penalize bad supervision and improve student model. Experiments on CIFAR-100, CINIC-10 and Tiny ImageNet show that our methods get encouraging performance compared with state-of-the-art methods. When combined with other KD-based methods, the performance will be further improved.

📄 PDF Abstract BibTeX arXiv:1911.07471

Code (1)

SforAiDl/KD_Lib pytorch

Tasks

Knowledge Distillation

Similar Papers 제목 키워드 기반

Learning the Wrong Lessons: Inserting Trojans During Knowledge Distillation

2023-03-09 · Leonard Tang, Tom Shlomi, Alexander Cai

In recent years, knowledge distillation has become a cornerstone of efficiently deployed machine learning, with labs and industries using knowledge distillation to train models that are inexpensive and resource-optimized…

Knowledge Distillation

FEED: Feature-level Ensemble for Knowledge Distillation

2019-09-24 · SeongUk Park, Nojun Kwak

Knowledge Distillation (KD) aims to transfer knowledge in a teacher-student framework, by providing the predictions of the teacher network to the student network in the training stage to help the student network generali…

Knowledge Distillation

GOLD: Generalized Knowledge Distillation via Out-of-Distribution-Guided Language Data Generation

2024-03-28 · Mohsen Gholami, Mohammad Akbari, Cindy Hu, Vaden Masrani 외

Knowledge distillation from LLMs is essential for the efficient deployment of language models. Prior works have proposed data generation using LLMs for preparing distilled models. We argue that generating data with LLMs …

Data-free Knowledge DistillationKnowledge Distillation

Time Series Classification: Lessons Learned in the (Literal) Field while Studying Chicken Behavior

2019-11-21 · Alireza Abdoli, Amy C. Murillo, Alec C. Gerry, Eamonn J. Keogh

Poultry farms are a major contributor to the human food chain. However, around the world, there have been growing concerns about the quality of life for the livestock in poultry farms; and increasingly vocal demands for …

ClusteringGeneral ClassificationTime SeriesTime Series Analysis+1

Preparing Lessons for Progressive Training on Language Models

2024-01-17 · Yu Pan, Ye Yuan, Yichun Yin, Jiaxin Shi 외

The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small mod…