paper-with-me

홈 › Papers

FCSN: Global Context Aware Segmentation by Learning the Fourier Coefficients of Objects in Medical Images

2022-07-29 · Young Seok Jeon, Hongfei Yang, Mengling Feng

The encoder-decoder model is a commonly used Deep Neural Network (DNN) model for medical image segmentation. Conventional encoder-decoder models make pixel-wise predictions focusing heavily on local patterns around the pixel. This makes it challenging to give segmentation that preserves the object's shape and topology, which often requires an understanding of the global context of the object. In this work, we propose a Fourier Coefficient Segmentation Network~(FCSN) -- a novel DNN-based model that segments an object by learning the complex Fourier coefficients of the object's masks. The Fourier coefficients are calculated by integrating over the whole contour. Therefore, for our model to make a precise estimation of the coefficients, the model is motivated to incorporate the global context of the object, leading to a more accurate segmentation of the object's shape. This global context awareness also makes our model robust to unseen local perturbations during inference, such as additive noise or motion blur that are prevalent in medical images. When FCSN is compared with other state-of-the-art models (UNet+, DeepLabV3+, UNETR) on 3 medical image segmentation tasks (ISIC\_2018, RIM\_CUP, RIM\_DISC), FCSN attains significantly lower Hausdorff scores of 19.14 (6\%), 17.42 (6\%), and 9.16 (14\%) on the 3 tasks, respectively. Moreover, FCSN is lightweight by discarding the decoder module, which incurs significant computational overhead. FCSN only requires 22.2M parameters, 82M and 10M fewer parameters than UNETR and DeepLabV3+. FCSN attains inference and training speeds of 1.6ms/img and 6.3ms/img, that is 8$\times$ and 3$\times$ faster than UNet and UNETR.

📄 PDF Abstract BibTeX arXiv:2207.14477

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

Ensemble learning in CNN augmented with fully connected subnetworks

2020-03-19 · Daiki Hirata, Norikazu Takahashi

Convolutional Neural Networks (CNNs) have shown remarkable performance in general object recognition tasks. In this paper, we propose a new model called EnsNet which is composed of one base CNN and multiple Fully Connect…

Ensemble LearningImage ClassificationObject Recognition

A Supervised Segmentation Network for Hyperspectral Image Classification

2021-02-04 · IEEE Transactions on Image Processing 2021 2 · Hao Sun, Xiangtao Zheng, Xiaoqiang Lu

Recently, deep learning has drawn broad attention in the hyperspectral image (HSI) classification task. Many works have focused on elaborately designing various spectral-spatial networks, where convolutional neural netwo…

ClassificationDiversityHyperspectral Image Classificationimage-classification+1

Face-Focused Cross-Stream Network for Deception Detection in Videos

2018-12-11 · CVPR 2019 6 · Mingyu Ding, An Zhao, Zhiwu Lu, Tao Xiang 외

Automated deception detection (ADD) from real-life videos is a challenging task. It specifically needs to address two problems: (1) Both face and body contain useful cues regarding whether a subject is deceptive. How to …

Deception DetectionDeception Detection In VideosEmotion RecognitionFace Detection+1

NPNet: A Non-Parametric Network with Adaptive Gaussian-Fourier Positional Encoding for 3D Classification and Segmentation

2026-01-31 · Mohammad Saeid, Amir Salarpour, Pedram MohajerAnsari, Mert D. Pesé arxiv

We present NPNet, a fully non-parametric approach for 3D point-cloud classification and part segmentation. NPNet contains no learned weights; instead, it builds point features using deterministic operators such as farthe…

3D Classification

FAN-Unet: Enhancing Unet with vision Fourier Analysis Block for Biomedical Image Segmentation

2024-11-28 · Jiashu Xu

Medical image segmentation is a critical aspect of modern medical research and clinical practice. Despite the remarkable performance of Convolutional Neural Networks (CNNs) in this domain, they inherently struggle to cap…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation