paper-with-me

홈 › Papers

Vision Xformers: Efficient Attention for Image Classification

2021-07-05 · Pranav Jeevan, Amit Sethi

Although transformers have become the neural architectures of choice for natural language processing, they require orders of magnitude more training data, GPU memory, and computations in order to compete with convolutional neural networks for computer vision. The attention mechanism of transformers scales quadratically with the length of the input sequence, and unrolled images have long sequence lengths. Plus, transformers lack an inductive bias that is appropriate for images. We tested three modifications to vision transformer (ViT) architectures that address these shortcomings. Firstly, we alleviate the quadratic bottleneck by using linear attention mechanisms, called X-formers (such that, X in {Performer, Linformer, Nystr\"omformer}), thereby creating Vision X-formers (ViXs). This resulted in up to a seven times reduction in the GPU memory requirement. We also compared their performance with FNet and multi-layer perceptron mixers, which further reduced the GPU memory requirement. Secondly, we introduced an inductive bias for images by replacing the initial linear embedding layer by convolutional layers in ViX, which significantly increased classification accuracy without increasing the model size. Thirdly, we replaced the learnable 1D position embeddings in ViT with Rotary Position Embedding (RoPE), which increases the classification accuracy for the same model size. We believe that incorporating such changes can democratize transformers by making them accessible to those with limited data and computing resources.

📄 PDF Abstract BibTeX arXiv:2107.02239

Code (2)

pranavphoenix/ViX 공식 구현 pytorch
pranavphoenix/VisionXformer pytorch

Tasks

ClassificationGPUimage-classificationImage ClassificationInductive BiasPosition

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
FAVOR+ 설명 없음
Performer Performer is a Transformer architecture which can estimate regular…
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…
Nyströmformer 설명 없음

Similar Papers 제목 키워드 기반

Convolutional Xformers for Vision

2022-01-25 · Pranav Jeevan, Amit Sethi

Vision transformers (ViTs) have found only limited practical use in processing images, in spite of their state-of-the-art accuracy on certain benchmarks. The reason for their limited use include their need for larger tra…

GPUimage-classificationImage Classification

RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings

2026-05-11 · Byeongchan Kim, Arijit Sehanobish, Avinava Dubey, Min-hwan Oh 외 arxiv

We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions $f$. They lead to the new class of 3D-Transform…

Point Clouds

SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration

2024-10-03 · Jintao Zhang, Jia Wei, Haofeng Huang, Pengle Zhang 외

The transformer architecture predominates across various models. As the heart of the transformer, attention has a computational complexity of O(N^2), compared to O(N) for linear transformations. When handling large seque…

Image GenerationQuantizationVideo Generation

SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization

2024-11-17 · Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei 외

Although quantization for linear layers has been widely used, its application to accelerate the attention process remains limited. To further enhance the efficiency of attention computation compared to SageAttention whil…

Image GenerationQuantizationVideo Generation

Simple Local Attentions Remain Competitive for Long-Context Tasks

2021-12-14 · NAACL 2022 7 · Wenhan Xiong, Barlas Oğuz, Anchit Gupta, Xilun Chen 외

Many NLP tasks require processing long contexts beyond the length limit of pretrained models. In order to scale these models to longer text sequences, many efficient long-range attention variants have been proposed. Desp…