Convolution-Free Medical Image Segmentation using Transformers
Like other applications in computer vision, medical image segmentation has been most successfully addressed using deep learning models that rely on the convolution operation as their main building block. Convolutions enjoy important properties such as sparse interactions, weight sharing, and translation equivariance. These properties give convolutional neural networks (CNNs) a strong and useful inductive bias for vision tasks. In this work we show that a different method, based entirely on self-attention between neighboring image patches and without any convolution operations, can achieve competitive or better results. Given a 3D image block, our network divides it into $n^3$ 3D patches, where $n=3 \text{ or } 5$ and computes a 1D embedding for each patch. The network predicts the segmentation map for the center patch of the block based on the self-attention between these patch embeddings. We show that the proposed model can achieve segmentation accuracies that are better than the state of the art CNNs on three datasets. We also propose methods for pre-training this model on large corpora of unlabeled images. Our experiments show that with pre-training the advantage of our proposed network over CNNs can be significant when labeled training data is small.
Code (1)
Tasks
Image SegmentationInductive BiasMedical Image SegmentationSegmentationSemantic SegmentationTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ConvTransSeg: A Multi-resolution Convolution-Transformer Network for Medical Image Segmentation
Convolutional neural networks (CNNs) achieved the state-of-the-art performance in medical image segmentation due to their ability to extract highly complex feature representations. However, it is argued in recent studies…
DecoderImage SegmentationMedical Image SegmentationSegmentation+1TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
Medical image segmentation is an essential prerequisite for developing healthcare systems, especially for disease diagnosis and treatment planning. On various medical image segmentation tasks, the u-shaped architecture, …
Cardiac SegmentationDecoderImage SegmentationMedical Image Segmentation+3CiT-Net: Convolutional Neural Networks Hand in Hand with Vision Transformers for Medical Image Segmentation
The hybrid architecture of convolutional neural networks (CNNs) and Transformer are very popular for medical image segmentation. However, it suffers from two challenges. First, although a CNNs branch can capture the loca…
Image SegmentationMedical Image SegmentationSegmentationSemantic SegmentationTransformer-CNN Fused Architecture for Enhanced Skin Lesion Segmentation
The segmentation of medical images is important for the improvement and creation of healthcare systems, particularly for early disease detection and treatment planning. In recent years, the use of convolutional neural ne…
Image SegmentationLesion SegmentationMedical Image SegmentationSegmentation+2ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation
Transformers have been extensively studied in medical image segmentation to build pairwise long-range dependence. Yet, relatively limited well-annotated medical image data makes transformers struggle to extract diverse g…
Image SegmentationMedical Image SegmentationSemantic Segmentation