paper-with-me

홈 › Papers

Spatial Context-Aware Self-Attention Model For Multi-Organ Segmentation

2020-12-16 · Hao Tang, Xingwei Liu, Kun Han, Shanlin Sun, Narisu Bai, Xuming Chen, Huang Qian, Yong liu, Xiaohui Xie

Multi-organ segmentation is one of most successful applications of deep learning in medical image analysis. Deep convolutional neural nets (CNNs) have shown great promise in achieving clinically applicable image segmentation performance on CT or MRI images. State-of-the-art CNN segmentation models apply either 2D or 3D convolutions on input images, with pros and cons associated with each method: 2D convolution is fast, less memory-intensive but inadequate for extracting 3D contextual information from volumetric images, while the opposite is true for 3D convolution. To fit a 3D CNN model on CT or MRI images on commodity GPUs, one usually has to either downsample input images or use cropped local regions as inputs, which limits the utility of 3D models for multi-organ segmentation. In this work, we propose a new framework for combining 3D and 2D models, in which the segmentation is realized through high-resolution 2D convolutions, but guided by spatial contextual information extracted from a low-resolution 3D model. We implement a self-attention mechanism to control which 3D features should be used to guide 2D segmentation. Our model is light on memory usage but fully equipped to take 3D contextual information into account. Experiments on multiple organ segmentation datasets demonstrate that by taking advantage of both 2D and 3D models, our method is consistently outperforms existing 2D and 3D models in organ segmentation accuracy, while being able to directly take raw whole-volume image data as inputs.

📄 PDF Abstract BibTeX arXiv:2012.09279

Code (0)

등록된 구현이 없습니다.

Tasks

Image SegmentationMedical Image AnalysisOrgan SegmentationSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

3D CNN 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Spatially Aware Multimodal Transformers for TextVQA

2020-07-23 · ECCV 2020 8 · Yash Kant, Dhruv Batra, Peter Anderson, Alex Schwing 외

Textual cues are essential for everyday tasks like buying groceries and using public transport. To develop this assistive technology, we study the TextVQA task, i.e., reasoning about text in images to answer a question. …

Optical Character Recognition (OCR)Spatial ReasoningTextVQAVisual Grounding+1

STJLA: A Multi-Context Aware Spatio-Temporal Joint Linear Attention Network for Traffic Forecasting

2021-12-04 · Yuchen Fang, Yanjun Qin, Haiyong Luo, Fang Zhao 외

Traffic prediction has gradually attracted the attention of researchers because of the increase in traffic big data. Therefore, how to mine the complex spatio-temporal correlations in traffic data to predict traffic cond…

PositionTime Series AnalysisTraffic Prediction

Spatial-Temporal Attention Network for Open-Set Fine-Grained Image Recognition

2022-11-25 · Jiayin Sun, Hong Wang, Qiulei Dong

Triggered by the success of transformers in various visual tasks, the spatial self-attention mechanism has recently attracted more and more attention in the computer vision community. However, we empirically found that a…

Fine-Grained Image RecognitionOpen Set Learning

CAViT -- Channel-Aware Vision Transformer for Dynamic Feature Fusion

2026-02-05 · Aon Safdar, Mohamed Saadeldin arxiv

Vision Transformers (ViTs) have demonstrated strong performance across a range of computer vision tasks by modeling long-range spatial interactions via self-attention. However, channel-wise mixing in ViTs remains static,…

Deep Multiple Instance Learning with Distance-Aware Self-Attention

2023-05-17 · Georg Wölflein, Lucie Charlotte Magister, Pietro Liò, David J. Harrison 외

Traditional supervised learning tasks require a label for every instance in the training set, but in many real-world applications, labels are only available for collections (bags) of instances. This problem setting, know…

Cancer Metastasis DetectionMultiple Instance LearningPosition