paper-with-me

홈 › Papers

ConvShareViT: Enhancing Vision Transformers with Convolutional Attention Mechanisms for Free-Space Optical Accelerators

2025-04-15 · Riad Ibadulla, Thomas M. Chen, Constantino Carlos Reyes-Aldasoro

This paper introduces ConvShareViT, a novel deep learning architecture that adapts Vision Transformers (ViTs) to the 4f free-space optical system. ConvShareViT replaces linear layers in multi-head self-attention (MHSA) and Multilayer Perceptrons (MLPs) with a depthwise convolutional layer with shared weights across input channels. Through the development of ConvShareViT, the behaviour of convolutions within MHSA and their effectiveness in learning the attention mechanism were analysed systematically. Experimental results demonstrate that certain configurations, particularly those using valid-padded shared convolutions, can successfully learn attention, achieving comparable attention scores to those obtained with standard ViTs. However, other configurations, such as those using same-padded convolutions, show limitations in attention learning and operate like regular CNNs rather than transformer models. ConvShareViT architectures are specifically optimised for the 4f optical system, which takes advantage of the parallelism and high-resolution capabilities of optical systems. Results demonstrate that ConvShareViT can theoretically achieve up to 3.04 times faster inference than GPU-based systems. This potential acceleration makes ConvShareViT an attractive candidate for future optical deep learning applications and proves that our ViT (ConvShareViT) can be employed using only the convolution operation, via the necessary optimisation of the ViT to balance performance and complexity.

📄 PDF Abstract BibTeX arXiv:2504.11517

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages

2025-04-21 · Zhoujie Qian

Vision Transformers (ViTs) have revolutionized computer vision by leveraging self-attention to model long-range dependencies. However, ViTs face challenges such as high computational costs due to the quadratic scaling of…

image-classificationImage Classification

Enhancing compact convolutional transformers with super attention

2025-08-26 · Simpenzwe Honore Leandre, Natenaile Asmamaw Shiferaw, Dillip Rout arxiv

In this paper, we propose a vision model that adopts token mixing, sequence-pooling, and convolutional tokenizers to achieve state-of-the-art performance and efficient inference in fixed context-length tasks. In the CIFA…

Data Augmentation

Optimizing Vision Transformers for Medical Image Segmentation

2022-10-14 · Qianying Liu, Chaitanya Kaul, Jun Wang, Christos Anagnostopoulos 외

For medical image semantic segmentation (MISS), Vision Transformers have emerged as strong alternatives to convolutional neural networks thanks to their inherent ability to capture long-range correlations. However, exist…

Domain AdaptationImage SegmentationMedical Image SegmentationSemantic Segmentation

CoSwin: Convolution Enhanced Hierarchical Shifted Window Attention For Small-Scale Vision

2025-09-10 · Puskal Khadka, Rodrigue Rizk, Longwei Wang, KC Santosh arxiv

Vision Transformers (ViTs) have achieved impressive results in computer vision by leveraging self-attention to model long-range dependencies. However, their emphasis on global context often comes at the expense of local …

Image Classification

DMFormer: Closing the Gap Between CNN and Vision Transformers

2022-09-16 · Zimian Wei, Hengyue Pan, Lujun Li, Menglong Lu 외

Vision transformers have shown excellent performance in computer vision tasks. As the computation cost of their self-attention mechanism is expensive, recent works tried to replace the self-attention mechanism in vision …

Inductive Biasobject-detectionObject DetectionSemantic Segmentation