Online Knowledge Distillation for Efficient Pose Estimation
Existing state-of-the-art human pose estimation methods require heavy computational resources for accurate predictions. One promising technique to obtain an accurate yet lightweight pose estimator is knowledge distillation, which distills the pose knowledge from a powerful teacher model to a less-parameterized student model. However, existing pose distillation works rely on a heavy pre-trained estimator to perform knowledge transfer and require a complex two-stage learning procedure. In this work, we investigate a novel Online Knowledge Distillation framework by distilling Human Pose structure knowledge in a one-stage manner to guarantee the distillation efficiency, termed OKDHP. Specifically, OKDHP trains a single multi-branch network and acquires the predicted heatmaps from each, which are then assembled by a Feature Aggregation Unit (FAU) as the target heatmaps to teach each branch in reverse. Instead of simply averaging the heatmaps, FAU which consists of multiple parallel transformations with different receptive fields, leverages the multi-scale information, thus obtains target heatmaps with higher-quality. Specifically, the pixel-wise Kullback-Leibler (KL) divergence is utilized to minimize the discrepancy between the target heatmaps and the predicted ones, which enables the student network to learn the implicit keypoint relationship. Besides, an unbalanced OKDHP scheme is introduced to customize the student networks with different compression rates. The effectiveness of our approach is demonstrated by extensive experiments on two common benchmark datasets, MPII and COCO.
Code (1)
Tasks
Knowledge DistillationPose EstimationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
UKD: Debiasing Conversion Rate Estimation via Uncertainty-regularized Knowledge Distillation
In online advertising, conventional post-click conversion rate (CVR) estimation models are trained using clicked samples. However, during online serving the models need to estimate for all impression ads, leading to the …
Knowledge DistillationSelection biasDistilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
Camera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for t…
Knowledge Distillation3D Object DetectionAutonomous DrivingEfficient training of lightweight neural networks using Online Self-Acquired Knowledge Distillation
Knowledge Distillation has been established as a highly promising approach for training compact and faster models by transferring knowledge from heavyweight and powerful models. However, KD in its conventional version co…
Density EstimationKnowledge DistillationADU-Depth: Attention-based Distillation with Uncertainty Modeling for Depth Estimation
Monocular depth estimation is challenging due to its inherent ambiguity and ill-posed nature, yet it is quite important to many applications. While recent works achieve limited accuracy by designing increasingly complica…
3D geometryDepth EstimationDomain AdaptationKnowledge Distillation+2On the Query Strategies for Efficient Online Active Distillation
Deep Learning (DL) requires lots of time and data, resulting in high computational demands. Recently, researchers employ Active Learning (AL) and online distillation to enhance training efficiency and real-time model ada…
Active LearningContinual LearningKnowledge DistillationPose Estimation