paper-with-me

Papers

DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

2021-06-03 · NeurIPS 2021 12 · Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie zhou, Cho-Jui Hsieh

Attention is sparse in vision transformers. We observe the final prediction in vision transformers is only based on a subset of most informative tokens, which is sufficient for accurate image recognition. Based on this observation, we propose a dynamic token sparsification framework to prune redundant tokens progressively and dynamically based on the input. Specifically, we devise a lightweight prediction module to estimate the importance score of each token given the current features. The module is added to different layers to prune redundant tokens hierarchically. To optimize the prediction module in an end-to-end manner, we propose an attention masking strategy to differentiably prune a token by blocking its interactions with other tokens. Benefiting from the nature of self-attention, the unstructured sparse tokens are still hardware friendly, which makes our framework easy to achieve actual speed-up. By hierarchically pruning 66% of the input tokens, our method greatly reduces 31%~37% FLOPs and improves the throughput by over 40% while the drop of accuracy is within 0.5% for various vision transformers. Equipped with the dynamic token sparsification framework, DynamicViT models can achieve very competitive complexity/accuracy trade-offs compared to state-of-the-art CNNs and vision transformers on ImageNet. Code is available at https://github.com/raoyongming/DynamicViT

📄 PDF Abstract BibTeX arXiv:2106.02034

Code (2)

raoyongming/DynamicViT 공식 구현 pytorch
vision-sjtu/quadmamba pytorch

Tasks

BlockingEfficient ViTsImage Classification

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Dynamic Spatial Sparsification for Efficient Vision Transformers and Convolutional Neural Networks

2022-07-04 · Yongming Rao, Zuyan Liu, Wenliang Zhao, Jie zhou 외

In this paper, we present a new approach for model acceleration by exploiting spatial sparsity in visual data. We observe that the final prediction in vision Transformers is only based on a subset of the most informative…

SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers

2025-11-13 · Oded Schlesinger, Amirhossein Farzam, J. Matias Di Martino, Guillermo Sapiro arxiv

While Vision Transformers (ViT) have demonstrated remarkable performance across diverse tasks, their computational demands are substantial, scaling quadratically with the number of processed tokens. Compact attention rep…

Computational Efficiency

DeSparsify: Adversarial Attack Against Token Sparsification Mechanisms in Vision Transformers

2024-02-04 · Oryan Yehezkel, Alon Zolfi, Amit Baras, Yuval Elovici 외

Vision transformers have contributed greatly to advancements in the computer vision domain, demonstrating state-of-the-art performance in diverse tasks (e.g., image classification, object detection). However, their high …

Adversarial AttackGPUimage-classificationImage Classification+2

Making Vision Transformers Efficient from A Token Sparsification View

2023-03-15 · CVPR 2023 1 · Shuning Chang, Pichao Wang, Ming Lin, Fan Wang 외

The quadratic computational complexity to the number of tokens limits the practical applications of Vision Transformers (ViTs). Several works propose to prune redundant tokens to achieve efficient ViTs. However, these me…

Efficient ViTsimage-classificationImage ClassificationInstance Segmentation+4

SaiT: Sparse Vision Transformers through Adaptive Token Pruning

2022-10-11 · Ling Li, David Thorsley, Joseph Hassoun

While vision transformers have achieved impressive results, effectively and efficiently accelerating these models can further boost performances. In this work, we propose a dense/sparse training framework to obtain a uni…

Knowledge Distillation