paper-with-me

홈 › Papers

TinyViT: Fast Pretraining Distillation for Small Vision Transformers

2022-07-21 · Kan Wu, Jinnian Zhang, Houwen Peng, Mengchen Liu, Bin Xiao, Jianlong Fu, Lu Yuan

Vision transformer (ViT) recently has drawn great attention in computer vision due to its remarkable model capability. However, most prevailing ViT models suffer from huge number of parameters, restricting their applicability on devices with limited resources. To alleviate this issue, we propose TinyViT, a new family of tiny and efficient small vision transformers pretrained on large-scale datasets with our proposed fast distillation framework. The central idea is to transfer knowledge from large pretrained models to small ones, while enabling small models to get the dividends of massive pretraining data. More specifically, we apply distillation during pretraining for knowledge transfer. The logits of large teacher models are sparsified and stored in disk in advance to save the memory cost and computation overheads. The tiny student transformers are automatically scaled down from a large pretrained model with computation and parameter constraints. Comprehensive experiments demonstrate the efficacy of TinyViT. It achieves a top-1 accuracy of 84.8% on ImageNet-1k with only 21M parameters, being comparable to Swin-B pretrained on ImageNet-21k while using 4.2 times fewer parameters. Moreover, increasing image resolutions, TinyViT can reach 86.5% accuracy, being slightly better than Swin-L while using only 11% parameters. Last but not the least, we demonstrate a good transfer ability of TinyViT on various downstream tasks. Code and models are available at https://github.com/microsoft/Cream/tree/main/TinyViT.

📄 PDF Abstract BibTeX arXiv:2207.10666

Code (3)

microsoft/cream 공식 구현 pytorch
rwightman/pytorch-image-models pytorch
https://gitlab.com/birder/birder pytorch

Tasks

Image ClassificationKnowledge Distillation

Similar Papers 제목 키워드 기반

TinyViT-Batten: Few-Shot Vision Transformer with Explainable Attention for Early Batten-Disease Detection on Pediatric MRI

2025-10-06 · Khartik Uppalapati, Bora Yimenicioglu, Shakeel Abdulkareem, Adan Eftekhari 외 arxiv

Batten disease (neuronal ceroid lipofuscinosis) is a rare pediatric neurodegenerative disorder whose early MRI signs are subtle and often missed. We propose TinyViT-Batten, a few-shot Vision Transformer (ViT) framework t…

Few-Shot Learning

RepViT-SAM: Towards Real-Time Segmenting Anything

2023-12-10 · Ao Wang, Hui Chen, Zijia Lin, Jungong Han 외

Segment Anything Model (SAM) has shown impressive zero-shot transfer performance for various computer vision tasks recently. However, its heavy computation costs remain daunting for practical applications. MobileSAM prop…

EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3

2025-11-19 · Chengxi Zeng, Yuxuan Jiang, Aaron Zhang arxiv

The Segment Anything Model 3 (SAM3) advances visual understanding with Promptable Concept Segmentation (PCS) across images and videos, but its unified architecture (shared vision backbone, DETR-style detector, dense-memo…

Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models

2023-11-30 · Raviteja Vemulapalli, Hadi Pouransari, Fartash Faghri, Sachin Mehta 외

Vision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, t…

Image RetrievalRetrievalTransfer Learning

Knowledge Distillation as Self-Supervised Learning

2022-01-17 · ICLR Track Blog 2022 5 · Anonymous

Self-supervised learning (SSL) methods have been shown to effectively train large neural networks with unlabeled data. These networks can produce useful image representations that can exceed the performance of supervised…

Knowledge DistillationSelf-Supervised Learning