paper-with-me

Papers

Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI

2026-04-06 · Mingjie Li, Edward Kim, Yue Zhao, Ehsan Adeli, Kilian M. Pohl arxiv

Learning a robust Variational Autoencoder (VAE) is a fundamental step for many deep learning applications in medical image analysis, such as MRI synthesizes. Existing brain VAEs predominantly focus on single-modality data (i.e., T1-weighted MRI), overlooking the complementary diagnostic value of other modalities like T2-weighted MRIs. Here, we propose a modality-aware and anatomically grounded 3D vector-quantized VAE (VQ-VAE) for reconstructing multi-modal brain MRIs. Called NeuroQuant, it first learns a shared latent representation across modalities using factorized multi-axis attention, which can capture relationships between distant brain regions. It then employs a dual-stream 3D encoder that explicitly separates the encoding of modality-invariant anatomical structures from modality-dependent appearance. Next, the anatomical encoding is discretized using a shared codebook and combined with modality-specific appearance features via Feature-wise Linear Modulation (FiLM) during the decoding phase. This entire approach is trained using a joint 2D/3D strategy in order to account for the slice-based acquisition of 3D MRI data. Extensive experiments on two multi-modal brain MRI datasets demonstrate that NeuroQuant achieves superior reconstruction fidelity compared to existing VAEs, enabling a scalable foundation for downstream generative modeling and cross-modal brain image analysis.

📄 PDF Abstract BibTeX arXiv:2604.05171

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes

2023-12-31 · Yuhta Takida, Yukara Ikemiya, Takashi Shibuya, Kazuki Shimada 외

Vector quantization (VQ) is a technique to deterministically learn features with discrete codebook representations. It is commonly performed with a variational autoencoding model, VQ-VAE, which can be further extended to…

QuantizationRepresentation Learning

General Point Model with Autoencoding and Autoregressive

2023-10-25 · Zhe Li, Zhangyang Gao, Cheng Tan, Stan Z. Li 외

The pre-training architectures of large language models encompass various types, including autoencoding models, autoregressive models, and encoder-decoder models. We posit that any modality can potentially benefit from a…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3

General Point Model Pretraining with Autoencoding and Autoregressive

2024-01-01 · CVPR 2024 1 · Zhe Li, Zhangyang Gao, Cheng Tan, Bocheng Ren 외

The pre-training architectures of large language models encompass various types including autoencoding models autoregressive models and encoder-decoder models. We posit that any modality can potentially benefit from …

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3

Topology-Constrained Quantized nnUNet for Efficient and Anatomically Accurate 3D Tooth Segmentation

2026-05-05 · Paarth Prasad, Ruchika Malhotra arxiv

We propose a topology-constrained quantized nnUNet framework for efficient and anatomically accurate 3D tooth segmentation, addressing the challenges of spatial distortion introduced by quantization in deep learning mode…

Medical Image SegmentationComputational Efficiency

VAMAE: Vessel-Aware Masked Autoencoders for OCT Angiography

2026-04-08 · Ilerioluwakiiye Abolade, Prince Mireku, Kelechi Chibundu, Peace Ododo 외 arxiv

Optical coherence tomography angiography (OCTA) provides non-invasive visualization of retinal microvasculature, but learning robust representations remains challenging due to sparse vessel structures and strong topologi…

Self-Supervised Learning