paper-with-me

홈 › Papers

Quality-Aware Modulation for Diffusion Transformers

2026-06-29 · Luke Budny, Yuhong Guo, Kevin Cheung arxiv

Modern text-to-image diffusion models, such as diffusion transformers (DiT), rely on timestep or prompt embeddings to modulate the strength of the denoising process in each timestep. While this modulation communicates the current noise level, it does not provide any quality-aware information, which can lead to generated images that are unaligned, visually inconsistent, and lacking in fidelity. In this paper, we propose the Quality Representation Module (QRM), a lightweight transformer module that learns a quality-aware representation based on existing model inputs, and produces a set of vectors $M_{qrm}$. These vectors adjust the adaptive LayerNorm modulation within the DiT transformer blocks, thereby injecting a quality-sensitive signal into the denoising parameters. The QRM introduces no significant changes to the sampling schedule or diffusion backbone. Experiments include ablations on QRM training losses and architectures, as well as empirical results demonstrating consistent image quality improvements over baseline DiT-based models.

📄 PDF Abstract BibTeX arXiv:2606.30934

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking Global Text Conditioning in Diffusion Transformers

2026-02-09 · Nikita Starodubcev, Daniil Pakhomov, Zongze Wu, Ilya Drobyshevskiy 외 arxiv

Diffusion transformers typically incorporate textual information via attention layers and a modulation mechanism using a pooled text embedding. Nevertheless, recent approaches discard modulation-based text conditioning a…

Video GenerationImage Editing

DP-aware AdaLN-Zero: Taming Conditioning-Induced Heavy-Tailed Gradients in Differentially Private Diffusion

2026-02-26 · Tao Huang, Jiayang Meng, Xu Yang, Chen Hou 외 arxiv

Condition injection enables diffusion models to generate context-aware outputs, which is essential for many time-series tasks. However, heterogeneous conditional contexts (e.g., observed history, missingness patterns or …

FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection

2026-01-02 · Ruiqiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang 외 arxiv

Large-scale text-to-image (T2I) diffusion models excel at open-domain synthesis but still struggle with precise text rendering, especially for multi-line layouts, dense typography, and long-tailed scripts such as Chinese…

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

2026-07-03 · Chaofan Gan, Zicheng Zhao, Yuanpeng Tu, Xi Chen 외 arxiv

Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiTs) remain insufficiently understood. In this work, we systematically…

Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers

2025-05-25 · Eric Tillman Bill, Cristian Perez Jensen, Sotiris Anagnostidis, Dimitri von Rütte

Denoising diffusion models exhibit remarkable generative capabilities, but remain challenging to train due to their inherent stochasticity, where high-variance gradient estimates lead to slow convergence. Previous works …

Denoising