paper-with-me

홈 › Papers

Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework

2021-10-04 · NeurIPS 2021 12 · Tengteng Huang, Yifan Sun, Xun Wang, Haotian Yao, Chi Zhang

Model smoothing is of central importance for obtaining a reliable teacher model in the student-teacher framework, where the teacher generates surrogate supervision signals to train the student. A popular model smoothing method is the Temporal Moving Average (TMA), which continuously averages the teacher parameters with the up-to-date student parameters. In this paper, we propose "Spatial Ensemble", a novel model smoothing mechanism in parallel with TMA. Spatial Ensemble randomly picks up a small fragment of the student model to directly replace the corresponding fragment of the teacher model. Consequentially, it stitches different fragments of historical student models into a unity, yielding the "Spatial Ensemble" effect. Spatial Ensemble obtains comparable student-teacher learning performance by itself and demonstrates valuable complementarity with temporal moving average. Their integration, named Spatial-Temporal Smoothing, brings general (sometimes significant) improvement to the student-teacher learning framework on a variety of state-of-the-art methods. For example, based on the self-supervised method BYOL, it yields +0.9% top-1 accuracy improvement on ImageNet, while based on the semi-supervised approach FixMatch, it increases the top-1 accuracy by around +6% on CIFAR-10 when only few training labels are available. Codes and models are available at: https://github.com/tengteng95/Spatial_Ensemble.

📄 PDF Abstract BibTeX arXiv:2110.01253

Code (1)

tengteng95/spatial_ensemble 공식 구현 pytorch

Tasks

Unity

Methods 이 논문이 사용한 방법론

FixMatch FixMatch is an algorithm that first generates pseudo-labels using the model's predictions on weakly-augmented unlabeled images. For a given image, the pseudo-label is only…
BYOL 설명 없음

Similar Papers 제목 키워드 기반

Why does Knowledge Distillation Work? Rethink its Attention and Fidelity Mechanism

2024-04-30 · Chenqi Guo, Shiwei Zhong, Xiaofeng Liu, Qianli Feng 외

Does Knowledge Distillation (KD) really work? Conventional wisdom viewed it as a knowledge transfer procedure where a perfect mimicry of the student to its teacher is desired. However, paradoxical studies indicate that c…

Data AugmentationDiversityKnowledge DistillationTransfer Learning

Revisiting Knowledge Distillation via Label Smoothing Regularization

2019-09-25 · CVPR 2020 6 · Li Yuan, Francis E. H. Tay, Guilin Li, Tao Wang 외

Knowledge Distillation (KD) aims to distill the knowledge of a cumbersome teacher model into a lightweight student model. Its success is generally attributed to the privileged information on similarities among categories…

Knowledge DistillationSelf-Knowledge Distillation

Consistently Informative Soft-Label Temperature for Knowledge Distillation

2026-05-19 · Hoang-Chau Luong, Nghia Van Vo, Kaiqi Zhao, Lingwei Chen arxiv

Knowledge distillation (KD) transfers knowledge from a high-capacity teacher to a compact student by matching their predictive distributions, with temperature scaling serving as a central mechanism for smoothing teacher …

Knowledge Distillation

An Ensemble Teacher-Student Learning Approach with Poisson Sub-sampling to Differential Privacy Preserving Speech Recognition

2022-10-12 · Chao-Han Huck Yang, Jun Qi, Sabato Marco Siniscalchi, Chin-Hui Lee

We propose an ensemble learning framework with Poisson sub-sampling to effectively train a collection of teacher models to issue some differential privacy (DP) guarantee for training data. Through boosting under DP, a st…

Ensemble LearningPrivacy Preservingspeech-recognitionSpeech Recognition+1

Improved Feature Distillation via Projector Ensemble

2022-10-27 · Yudong Chen, Sen Wang, Jiajun Liu, Xuwei Xu 외

In knowledge distillation, previous feature distillation methods mainly focus on the design of loss functions and the selection of the distilled layers, while the effect of the feature projector between the student and t…

Knowledge DistillationMulti-Task Learning