paper-with-me

Papers

VolDiT: Controllable Volumetric Medical Image Synthesis with Diffusion Transformers

2026-03-26 · Marvin Seyfarth, Salman Ul Hassan Dar, Yannik Frisch, Philipp Wild, Norbert Frey, Florian André, Sandy Engelhardt arxiv

Diffusion models have become a leading approach for high-fidelity medical image synthesis. However, most existing methods for 3D medical image generation rely on convolutional U-Net backbones within latent diffusion frameworks. While effective, these architectures impose strong locality biases and limited receptive fields, which may constrain scalability, global context integration, and flexible conditioning. In this work, we introduce VolDiT, the first purely transformer-based 3D Diffusion Transformer for volumetric medical image synthesis. Our approach extends diffusion transformers to native 3D data through volumetric patch embeddings and global self-attention operating directly over 3D tokens. To enable structured control, we propose a timestep-gated control adapter that maps segmentation masks into learnable control tokens that modulate transformer layers during denoising. This token-level conditioning mechanism allows precise spatial guidance while preserving the modeling advantages of transformer architectures. We evaluate our model on high-resolution 3D medical image synthesis tasks and compare it to state-of-the-art 3D latent diffusion models based on U-Nets. Results demonstrate improved global coherence, superior generative fidelity, and enhanced controllability. Our findings suggest that fully transformerbased diffusion models provide a flexible foundation for volumetric medical image synthesis. The code and models trained on public data are available at https://github.com/Cardio-AI/voldit.

📄 PDF Abstract BibTeX arXiv:2603.25181

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Image Generation

Similar Papers 제목 키워드 기반

Make-A-Volume: Leveraging Latent Diffusion Models for Cross-Modality 3D Brain MRI Synthesis

2023-07-19 · Lingting Zhu, Zeyue Xue, Zhenchao Jin, Xian Liu 외

Cross-modality medical image synthesis is a critical topic and has the potential to facilitate numerous applications in the medical imaging field. Despite recent successes in deep-learning-based generative models, most c…

Computational EfficiencyImage Generation

SAINT: Spatially Aware Interpolation NeTwork for Medical Slice Synthesis

2020-06-01 · CVPR 2020 6 · Cheng Peng, Wei-An Lin, Haofu Liao, Rama Chellappa 외

Deep learning-based single image super-resolution (SISR) methods face various challenges when applied to 3D medical volumetric data (i.e., CT and MR images) due to the high memory cost and anisotropic resolution, which a…

Image Super-ResolutionSuper-Resolution

CoNFies: Controllable Neural Face Avatars

2022-11-16 · Heng Yu, Koichiro Niinuma, Laszlo A. Jeni

Neural Radiance Fields (NeRF) are compelling techniques for modeling dynamic 3D scenes from 2D image collections. These volumetric representations would be well suited for synthesizing novel facial expressions but for tw…

Action RecognitionNeRF

Explicit Temporal Embedding in Deep Generative Latent Models for Longitudinal Medical Image Synthesis

2023-01-13 · Julian Schön, Raghavendra Selvan, Lotte Nygård, Ivan Richter Vogelius 외

Medical imaging plays a vital role in modern diagnostics and treatment. The temporal nature of disease or treatment progression often results in longitudinal data. Due to the cost and potential harm, acquiring large medi…

DisentanglementImage Generation

FLAME-in-NeRF : Neural control of Radiance Fields for Free View Face Animation

2021-08-10 · ShahRukh Athar, Zhixin Shu, Dimitris Samaras

This paper presents a neural rendering method for controllable portrait video synthesis. Recent advances in volumetric neural rendering, such as neural radiance fields (NeRF), has enabled the photorealistic novel view sy…

Face ModelNeRFNeural RenderingNovel View Synthesis