paper-with-me

홈 › Papers

Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation

2026-06-18 · Zhenkai Zhang, Markus Hiller, Krista A. Ehinger, Tom Drummond arxiv

Generating high-resolution 3D CT volumes with fine details remains challenging due to substantial computational demands and optimization difficulties inherent to existing generative models. In this paper, we propose the Pixel-Level Residual Diffusion Transformer (PRDiT), a scalable generative framework that synthesizes high-quality 3D medical volumes directly at voxel-level. PRDiT introduces a two-stage training architecture comprising 1) a local denoiser in the form of an MLP-based blind estimator operating on overlapping 3D patches to separate low-frequency structures efficiently, and 2) a global residual diffusion transformer employing memory-efficient attention to model and refine high-frequency residuals across entire volumes. This coarse-to-fine modeling strategy simplifies optimization, enhances training stability, and effectively preserves subtle structures without the limitations of an autoencoder bottleneck. Extensive experiments conducted on the LIDC-IDRI and RAD-ChestCT datasets demonstrate that PRDiT consistently outperforms state-of-the-art models, such as HA-GAN, 3D LDM and WDM-3D, achieving significantly lower 3D FID, MMD and Wasserstein distance scores.

📄 PDF Abstract BibTeX arXiv:2606.20112

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers

2024-01-21 · Katherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham 외

We present the Hourglass Diffusion Transformer (HDiT), an image generative model that exhibits linear scaling with pixel count, supporting training at high-resolution (e.g. $1024 \times 1024$) directly in pixel-space. Bu…

Image Generation

PixelDiT: Pixel Diffusion Transformers for Image Generation

2025-11-25 · Yongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng 외 arxiv

Latent-space modeling has been the standard for Diffusion Transformers (DiTs). However, it relies on a two-stage pipeline where the pretrained autoencoder introduces lossy reconstruction, leading to error accumulation wh…

Text-to-Image Generation

Neural Residual Diffusion Models for Deep Scalable Vision Generation

2024-06-19 · Zhiyuan Ma, Liangliang Zhao, Biqing Qi, BoWen Zhou

The most advanced diffusion models have recently adopted increasingly deep stacked networks (e.g., U-Net or Transformer) to promote the generative emergence capabilities of vision generation models similar to large langu…

Denoising

Pixel-Space Diffusion Transformers

2026-07-20 · Renye Yan, Jikang Cheng, You Wu, Ling Liang 외 arxiv

Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while sepa…

Normality-Preserving Continual Industrial Anomaly Detection via Orthogonal LoRA Banks

2026-06-01 · Weibai Fang, Haijun Che, Feiyang Ren, Qiancheng Lao arxiv

Continual industrial anomaly detection with diffusion models suffers from historical normality prior drift and catastrophic forgetting. Existing continual diffusion methods preserve previous knowledge through replay or c…

Anomaly Detection