Co-training $2^L$ Submodels for Visual Recognition
We introduce submodel co-training, a regularization method related to co-training, self-distillation and stochastic depth. Given a neural network to be trained, for each sample we implicitly instantiate two altered networks, ``submodels'', with stochastic depth: we activate only a subset of the layers. Each network serves as a soft teacher to the other, by providing a loss that complements the regular loss provided by the one-hot label. Our approach, dubbed cosub, uses a single set of weights, and does not involve a pre-trained external model or temporal averaging. Experimentally, we show that submodel co-training is effective to train backbones for recognition tasks such as image classification and semantic segmentation. Our approach is compatible with multiple architectures, including RegNet, ViT, PiT, XCiT, Swin and ConvNext. Our training strategy improves their results in comparable settings. For instance, a ViT-B pretrained with cosub on ImageNet-21k obtains 87.4% top-1 acc. @448 on ImageNet-val.
Code (1)
Tasks
image-classificationImage ClassificationSemantic SegmentationSimilar Papers 제목 키워드 기반
Co-Training 2L Submodels for Visual Recognition
This paper introduces submodel co-training, a regularization method related to co-training, self-distillation and stochastic depth. Given a neural network to be trained, for each sample we implicitly instantiate two …
image-classificationImage ClassificationSemantic SegmentationDropout Inference with Non-Uniform Weight Scaling
Dropout as regularization has been used extensively to prevent overfitting for training neural networks. During training, units and their connections are randomly dropped, which could be considered as sampling many diffe…
Sharing pattern submodels for prediction with missing values
Missing values are unavoidable in many applications of machine learning and present challenges both during training and at test time. When variables are missing in recurring patterns, fitting separate pattern submodels h…
ImputationMissing ValuesPredictionA Scalable Model Specialization Framework for Training and Inference using Submodels and its Application to Speech Model Personalization
Model fine-tuning and adaptation have become a common approach for model specialization for downstream tasks or domains. Fine-tuning the entire model or a subset of the parameters using light-weight adaptation has shown …
modelMeta-Learning-Based Deep Reinforcement Learning for Multiobjective Optimization Problems
Deep reinforcement learning (DRL) has recently shown its success in tackling complex combinatorial optimization problems. When these problems are extended to multiobjective ones, it becomes difficult for the existing DRL…
Combinatorial OptimizationDeep Reinforcement LearningDiversityMeta-Learning+4