paper-with-me

홈 › Papers

Knowledge Distillation $\approx$ Label Smoothing: Fact or Fallacy?

2023-01-30 · Md Arafat Sultan

Originally proposed as a method for knowledge transfer from one model to another, some recent studies have suggested that knowledge distillation (KD) is in fact a form of regularization. Perhaps the strongest argument of all for this new perspective comes from its apparent similarities with label smoothing (LS). Here we re-examine this stated equivalence between the two methods by comparing the predictive confidences of the models they train. Experiments on four text classification tasks involving models of different sizes show that: (a) In most settings, KD and LS drive model confidence in completely opposite directions, and (b) In KD, the student inherits not only its knowledge but also its confidence from the teacher, reinforcing the classical knowledge transfer view.

📄 PDF Abstract BibTeX arXiv:2301.12609

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillationtext-classificationText ClassificationTransfer Learning

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

Adaptive Label Smoothing with Self-Knowledge in Natural Language Generation

2022-10-22 · Dongkyu Lee, Ka Chun Cheung, Nevin L. Zhang

Overconfidence has been shown to impair generalization and calibration of a neural network. Previous studies remedy this issue by adding a regularization term to a loss function, preventing a model from making a peaked d…

Knowledge DistillationText Generation

Adaptive Label Smoothing with Self-Knowledge

2021-09-29 · Dongkyu Lee, Ka Chun Cheung, Nevin Zhang

Overconfidence has been shown to impair generalization and calibration of a neural network. Previous studies remedy this issue by adding a regularization term to a loss function, preventing a model from making a peaked d…

Knowledge DistillationMachine Translation

Pseudo Knowledge Distillation: Towards Learning Optimal Instance-specific Label Smoothing Regularization

2021-09-29 · Peng Lu, Ahmad Rashid, Ivan Kobyzev, Mehdi Rezagholizadeh 외

Knowledge Distillation (KD) is an algorithm that transfers the knowledge of a trained, typically larger, neural network into another model under training. Although a complete understanding of KD is elusive, a growing bod…

image-classificationImage ClassificationKnowledge DistillationNatural Language Understanding

Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study

2021-04-01 · ICLR 2021 1 · Zhiqiang Shen, Zechun Liu, Dejia Xu, Zitian Chen 외

This work aims to empirically clarify a recently discovered perspective that label smoothing is incompatible with knowledge distillation. We begin by introducing the motivation behind on how this incompatibility is raise…

image-classificationImage ClassificationKnowledge DistillationMachine Translation+1

Distilling Knowledge from Pre-trained Language Models via Text Smoothing

2020-05-08 · Xing Wu, Yibing Liu, Xiangyang Zhou, dianhai yu

This paper studies compressing pre-trained language models, like BERT (Devlin et al.,2019), via teacher-student knowledge distillation. Previous works usually force the student model to strictly mimic the smoothed labels…

Knowledge DistillationLanguage ModelingLanguage Modelling