Similarity-Preserving Knowledge Distillation
Knowledge distillation is a widely applicable technique for training a student neural network under the guidance of a trained teacher network. For example, in neural network compression, a high-capacity teacher is distilled to train a compact student; in privileged learning, a teacher trained with privileged data is distilled to train a student without access to that data. The distillation loss determines how a teacher's knowledge is captured and transferred to the student. In this paper, we propose a new form of knowledge distillation loss that is inspired by the observation that semantically similar inputs tend to elicit similar activation patterns in a trained network. Similarity-preserving knowledge distillation guides the training of a student network such that input pairs that produce similar (dissimilar) activations in the teacher network produce similar (dissimilar) activations in the student network. In contrast to previous distillation methods, the student is not required to mimic the representation space of the teacher, but rather to preserve the pairwise similarities in its own representation space. Experiments on three public datasets demonstrate the potential of our approach.
Code (1)
Tasks
Knowledge DistillationNeural Network CompressionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Novel Self-Knowledge Distillation Approach with Siamese Representation Learning for Action Recognition
Knowledge distillation is an effective transfer of knowledge from a heavy network (teacher) to a small network (student) to boost students' performance. Self-knowledge distillation, the special case of knowledge distilla…
Action RecognitionKnowledge DistillationRepresentation LearningSelf-Knowledge DistillationOn Representation Knowledge Distillation for Graph Neural Networks
Knowledge distillation is a learning paradigm for boosting resource-efficient graph neural networks (GNNs) using more expressive yet cumbersome teacher models. Past work on distillation for GNNs proposed the Local Struct…
Contrastive LearningKnowledge DistillationCategorical Relation-Preserving Contrastive Knowledge Distillation for Medical Image Classification
The amount of medical images for training deep classification models is typically very scarce, making these deep models prone to overfit the training data. Studies showed that knowledge distillation (KD), especially the …
Classificationimage-classificationImage ClassificationKnowledge Distillation+2Two-Step Knowledge Distillation for Tiny Speech Enhancement
Tiny, causal models are crucial for embedded audio machine learning applications. Model compression can be achieved via distilling knowledge from a large teacher into a smaller student model. In this work, we propose a n…
Knowledge DistillationModel CompressionSpeech EnhancementD3still: Decoupled Differential Distillation for Asymmetric Image Retrieval
Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However these one-to-one constraint approaches often fail to mai…
Image RetrievalRetrieval