paper-with-me

홈 › Papers

How Lightweight Can A Vision Transformer Be

2024-07-25 · Jen Hong Tan

In this paper, we explore a strategy that uses Mixture-of-Experts (MoE) to streamline, rather than augment, vision transformers. Each expert in an MoE layer is a SwiGLU feedforward network, where V and W2 are shared across the layer. No complex attention or convolutional mechanisms are employed. Depth-wise scaling is applied to progressively reduce the size of the hidden layer and the number of experts is increased in stages. Grouped query attention is used. We studied the proposed approach with and without pre-training on small datasets and investigated whether transfer learning works at this scale. We found that the architecture is competitive even at a size of 0.67M parameters.

📄 PDF Abstract BibTeX arXiv:2407.17783

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-ExpertsTransfer Learning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
MoE 설명 없음
SwiGLU SwiGLU is an activation function which is a variant of GLU. The definition is as follows: $$ \text{SwiGLU}\left(x, W, V, b, c,…

Similar Papers 제목 키워드 기반

Image Recognition with Online Lightweight Vision Transformer: A Survey

2025-05-06 · Zherui Zhang, Rongtao Xu, Jie zhou, Changwei Wang 외

The Transformer architecture has achieved significant success in natural language processing, motivating its adaptation to computer vision tasks. Unlike convolutional neural networks, vision transformers inherently captu…

Knowledge DistillationSurvey

Pre-training of Lightweight Vision Transformers on Small Datasets with Minimally Scaled Images

2024-02-06 · Jen Hong Tan

Can a lightweight Vision Transformer (ViT) match or exceed the performance of Convolutional Neural Networks (CNNs) like ResNet on small datasets with small image resolutions? This report demonstrates that a pure ViT can …

Image Classification

Rethinking Local Perception in Lightweight Vision Transformer

2023-03-31 · Qihang Fan, Huaibo Huang, Jiyang Guan, Ran He

Vision Transformers (ViTs) have been shown to be effective in various vision tasks. However, resizing them to a mobile-friendly size leads to significant performance degradation. Therefore, developing lightweight vision …

image-classificationImage Classificationobject-detectionObject Detection+1

Towards Lightweight Transformer via Group-wise Transformation for Vision-and-Language Tasks

2022-04-16 · Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Yan Wang 외

Despite the exciting performance, Transformer is criticized for its excessive parameters and computation cost. However, compressing Transformer remains as an open problem due to its internal complexity of the layer desig…

image-classificationImage Classification

Weed mapping in multispectral drone imagery using lightweight vision transformers

2023-12-28 · Neurocomputing 2023 12 · Giovanna Castellano, Pasquale De Marinis, Gennaro Vessio

In precision agriculture, non-invasive remote sensing can be used to observe crops and weeds in visible and non-visible spectra. This paper proposes a novel approach for weed mapping using lightweight Vision Transformers…

ManagementSemantic SegmentationTransfer Learning