paper-with-me

홈 › Papers

Distilling Knowledge by Mimicking Features

2020-11-03 · Guo-Hua Wang, Yifan Ge, Jianxin Wu

Knowledge distillation (KD) is a popular method to train efficient networks ("student") with the help of high-capacity networks ("teacher"). Traditional methods use the teacher's soft logits as extra supervision to train the student network. In this paper, we argue that it is more advantageous to make the student mimic the teacher's features in the penultimate layer. Not only the student can directly learn more effective information from the teacher feature, feature mimicking can also be applied for teachers trained without a softmax layer. Experiments show that it can achieve higher accuracy than traditional KD. To further facilitate feature mimicking, we decompose a feature vector into the magnitude and the direction. We argue that the teacher should give more freedom to the student feature's magnitude, and let the student pay more attention on mimicking the feature direction. To meet this requirement, we propose a loss term based on locality-sensitive hashing (LSH). With the help of this new loss, our method indeed mimics feature directions more accurately, relaxes constraints on feature magnitudes, and achieves state-of-the-art distillation accuracy. We provide theoretical analyses of how LSH facilitates feature direction mimicking, and further extend feature mimicking to multi-label recognition and object detection.

📄 PDF Abstract BibTeX arXiv:2011.01424

Code (3)

DoctorKey/LSHFM.detection 공식 구현 pytorch
DoctorKey/LSHFM.multiclassification 공식 구현 pytorch
DoctorKey/LSHFM.singleclassification 공식 구현 pytorch

Tasks

Knowledge Distillationobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Localization Distillation for Dense Object Detection

2021-02-24 · CVPR 2022 1 · Zhaohui Zheng, Rongguang Ye, Ping Wang, Dongwei Ren 외

Knowledge distillation (KD) has witnessed its powerful capability in learning compact models in object detection. Previous KD methods for object detection mostly focus on imitating deep features within the imitation regi…

Dense Object DetectionKnowledge DistillationObjectobject-detection+1

CrossKD: Cross-Head Knowledge Distillation for Object Detection

2023-06-20 · CVPR 2024 1 · Jiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li 외

Knowledge Distillation (KD) has been validated as an effective model compression technique for learning compact object detectors. Existing state-of-the-art KD methods for object detection are mostly based on feature imit…

Dense Object DetectionKnowledge DistillationModel CompressionObject+2

Cumulative Spatial Knowledge Distillation for Vision Transformers

2023-07-17 · ICCV 2023 1 · Borui Zhao, RenJie Song, Jiajun Liang

Distilling knowledge from convolutional neural networks (CNNs) is a double-edged sword for vision transformers (ViTs). It boosts the performance since the image-friendly local-inductive bias of CNN helps ViT learn faster…

Inductive BiasKnowledge DistillationTransfer Learning

Learning to Steer by Mimicking Features from Heterogeneous Auxiliary Networks

2018-11-07 · Yuenan Hou, Zheng Ma, Chunxiao Liu, Chen Change Loy

The training of many existing end-to-end steering angle prediction models heavily relies on steering angles as the supervisory signal. Without learning from much richer contexts, these methods are susceptible to the pres…

Image SegmentationMulti-Task LearningOptical Flow EstimationSemantic Segmentation+1

Cross-Architecture Knowledge Distillation

2022-07-12 · Yufan Liu, Jiajiong Cao, Bing Li, Weiming Hu 외

Transformer attracts much attention because of its ability to learn global relations and superior performance. In order to achieve higher performance, it is natural to distill complementary knowledge from Transformer to …

Knowledge Distillation