How Lightweight Can A Vision Transformer Be
In this paper, we explore a strategy that uses Mixture-of-Experts (MoE) to streamline, rather than augment, vision transformers. Each expert in an MoE layer is a SwiGLU feedforward network, where V and W2 are shared across the layer. No complex attention or convolutional mechanisms are employed. Depth-wise scaling is applied to progressively reduce the size of the hidden layer and the number of experts is increased in stages. Grouped query attention is used. We studied the proposed approach with and without pre-training on small datasets and investigated whether transfer learning works at this scale. We found that the architecture is competitive even at a size of 0.67M parameters.
Code (0)
등록된 구현이 없습니다.
Tasks
Mixture-of-ExpertsTransfer LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Image Recognition with Online Lightweight Vision Transformer: A Survey
The Transformer architecture has achieved significant success in natural language processing, motivating its adaptation to computer vision tasks. Unlike convolutional neural networks, vision transformers inherently captu…
Knowledge DistillationSurveyPre-training of Lightweight Vision Transformers on Small Datasets with Minimally Scaled Images
Can a lightweight Vision Transformer (ViT) match or exceed the performance of Convolutional Neural Networks (CNNs) like ResNet on small datasets with small image resolutions? This report demonstrates that a pure ViT can …
Image ClassificationRethinking Local Perception in Lightweight Vision Transformer
Vision Transformers (ViTs) have been shown to be effective in various vision tasks. However, resizing them to a mobile-friendly size leads to significant performance degradation. Therefore, developing lightweight vision …
image-classificationImage Classificationobject-detectionObject Detection+1Towards Lightweight Transformer via Group-wise Transformation for Vision-and-Language Tasks
Despite the exciting performance, Transformer is criticized for its excessive parameters and computation cost. However, compressing Transformer remains as an open problem due to its internal complexity of the layer desig…
image-classificationImage ClassificationWeed mapping in multispectral drone imagery using lightweight vision transformers
In precision agriculture, non-invasive remote sensing can be used to observe crops and weeds in visible and non-visible spectra. This paper proposes a novel approach for weed mapping using lightweight Vision Transformers…
ManagementSemantic SegmentationTransfer Learning