paper-with-me

홈 › Papers

GPAFormer: Graph-guided Patch Aggregation Transformer for Efficient 3D Medical Image Segmentation

2026-04-08 · Chung-Ming Lo, I-Yun Liu, Wei-Yang Lin arxiv

Deep learning has been widely applied to 3D medical image segmentation tasks. However, due to the diversity of imaging modalities, the high-dimensional nature of the data, and the heterogeneity of anatomical structures, achieving both segmentation accuracy and computational efficiency in multi-organ segmentation remains a challenge. This study proposed GPAFormer, a lightweight network architecture specifically designed for 3D medical image segmentation, emphasizing efficiency while keeping high accuracy. GPAFormer incorporated two core modules: the multi-scale attention-guided stacked aggregation (MASA) and the mutual-aware patch graph aggregator (MPGA). MASA utilized three parallel paths with different receptive fields, combined through planar aggregation, to enhance the network's capability in handling structures of varying sizes. MPGA employed a graph-guided approach to dynamically aggregate regions with similar feature distributions based on inter-patch feature similarity and spatial adjacency, thereby improving the discrimination of both internal and boundary structures of organs. Experiments were performed on public whole-body CT and MRI datasets including BTCV, Synapse, ACDC, and BraTS. Compared to the existed 3D segmentation networkd, GPAFormer using only 1.81 M parameters achieved overall highest DSC on BTCV (75.70%), Synapse (81.20%), ACDC (89.32%), and BraTS (82.74%). Using consumer level GPU, the inference time for one validation case of BTCV spent less than one second. The results demonstrated that GPAFormer balanced accuracy and efficiency in multi-organ, multi-modality 3D segmentation tasks across various clinical scenarios especially for resource-constrained and time-sensitive clinical environments.

📄 PDF Abstract BibTeX arXiv:2604.06658

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Image SegmentationComputational Efficiency

Similar Papers 제목 키워드 기반

Pose-guided Feature Disentangling for Occluded Person Re-identification Based on Transformer

2021-12-05 · Tao Wang, Hong Liu, Pinhao Song, Tianyu Guo 외

Occluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem …

DecoderOccluded Person Re-IdentificationPerson Re-Identification

Self-learned representation-guided latent diffusion model for breast cancer classification in deep ultraviolet whole surface images

2026-01-16 · Pouya Afshin, David Helminiak, Tianling Niu, Julie M. Jorns 외 arxiv

Breast-Conserving Surgery (BCS) requires precise intraoperative margin assessment to preserve healthy tissue. Deep Ultraviolet Fluorescence Scanning Microscopy (DUV-FSM) offers rapid, high-resolution surface imaging for …

Self-Supervised LearningCancer Classification

A Transformer-Based Adaptive Semantic Aggregation Method for UAV Visual Geo-Localization

2024-01-03 · Shishen Li, Cuiwei Liu, Huaijun Qiu, Zhaokui Li

This paper addresses the task of Unmanned Aerial Vehicles (UAV) visual geo-localization, which aims to match images of the same geographic target taken by different platforms, i.e., UAVs and satellites. In general, the k…

geo-localization

Survival Modeling from Whole Slide Images via Patch-Level Graph Clustering and Mixture Density Experts

2025-07-22 · Ardhendu Sekhar, Vasu Soni, Keshav Aske, Garima Jain 외 arxiv

We propose a modular framework for predicting cancer specific survival directly from whole slide pathology images (WSIs). The framework consists of four key stages designed to capture prognostic and morphological heterog…

Graph Clustering

Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era

2025-11-08 · Feng Lu, Tong Jin, Canming Ye, Yunpeng Liu 외 arxiv

Visual place recognition (VPR) is typically regarded as a specific image retrieval task, whose core lies in representing images as global descriptors. Over the past decade, dominant VPR methods (e.g., NetVLAD) have follo…

Visual Place RecognitionImage Retrieval