paper-with-me

홈 › Papers

Structured Residual Connectivity Matters for Diffusion Transformers

2026-09-27 · Yuhe Liu, Xinyin Ma, Gongfan Fang, Songhua Liu, Xinchao Wang hf

Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers as a monolithic state. In this work, we rethink residual connections in diffusion transformers and propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising. First, we conduct a systematic analysis of DiT's internal representation, revealing a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by this, we introduce a structured connectivity design that explicitly integrates local residual connections with long-range pathways. Instead of static skip connections or dense all-layer routing, our method enables each transformer block to selectively ``attend'' to critical earlier representations, dynamically retrieving spatial and semantic cues through direct, differentiable cross-depth paths. Experiments show that our adaptive connectivity leads to faster convergence, with up to 1.73times fewer training iterations, and significant gains in FID and visual quality with less than 0.1% additional parameters, further improving a strong REPA-XL/2 model from 5.9 to 4.34 FID without guidance and reaching 1.39 FID with classifier-free guidance. Our findings suggest that adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that incorporating structured information pathways provides a simple and effective direction for improving scalable generative models.

📄 PDF Abstract BibTeX arXiv:2609.33203

Code (0)

등록된 구현이 없습니다.

Tasks

Image Denoising

Similar Papers 제목 키워드 기반

FreqFormer: Hierarchical Frequency-Domain Attention with Adaptive Spectral Routing for Long-Sequence Video Diffusion Transformers

2026-04-14 · Haopeng Jin arxiv

Long-sequence video diffusion transformers hit a quadratic self-attention cost that dominates runtime and memory for very long token sequences. Most efficient attention methods use one approximation everywhere, yet video…

Graph Residual Noise Learner Network for Brain Connectivity Graph Prediction

2024-09-30 · Oytun Demirbilek, Tingying Peng, Alaa Bessadok

A morphological brain graph depicting a connectional fingerprint is of paramount importance for charting brain dysconnectivity patterns. Such data often has missing observations due to various reasons such as time-consum…

High Frequency Matters: Uncertainty Guided Image Compression with Wavelet Diffusion

2024-07-17 · Juan Song, Jiaxiang He, Lijie Yang, Mingtao Feng 외

Diffusion probabilistic models have recently achieved remarkable success in generating high-quality images. However, balancing high perceptual quality and low distortion remains challenging in image compression applicati…

DecoderImage CompressionPrediction

Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers

2025-10-24 · Dogyun Park, Moayed Haji-Ali, Yanyu Li, Willi Menapace 외 arxiv

Diffusion Transformers (DiTs) deliver state-of-the-art generative performance but their quadratic training cost with sequence length makes large-scale pretraining prohibitively expensive. Token dropping can reduce traini…

TrioPose: Native Triple-Stream Diffusion Transformers for Pose-Guided Text-to-Image Generation

2026-06-05 · Dian Gu, Zhengyi Yang arxiv

Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios. While existing UNet-based adapters struggle with long-range spatial dependencies, emerging…

Text-to-Image Generation