paper-with-me

홈 › Papers

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

2024-03-05 · Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, Robin Rombach

Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and noise in a straight line. Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales. Through a large-scale study, we demonstrate the superior performance of this approach compared to established diffusion formulations for high-resolution text-to-image synthesis. Additionally, we present a novel transformer-based architecture for text-to-image generation that uses separate weights for the two modalities and enables a bidirectional flow of information between image and text tokens, improving text comprehension, typography, and human preference ratings. We demonstrate that this architecture follows predictable scaling trends and correlates lower validation loss to improved text-to-image synthesis as measured by various metrics and human evaluations. Our largest models outperform state-of-the-art models, and we will make our experimental data, code, and model weights publicly available.

📄 PDF Abstract BibTeX arXiv:2403.03206

Code (2)

Karine-Huang/T2I-CompBench pytorch
hxixixh/adaflow pytorch

Tasks

Image GenerationReading ComprehensionText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

I-Max: Maximize the Resolution Potential of Pre-trained Rectified Flow Transformers with Projected Flow

2024-10-10 · Ruoyi Du, Dongyang Liu, Le Zhuo, Qin Qi 외

Rectified Flow Transformers (RFTs) offer superior training and inference efficiency, making them likely the most viable direction for scaling up diffusion models. However, progress in generation resolution has been relat…

2k

NAMI: Efficient Image Generation via Progressive Rectified Flow Transformers

2025-03-12 · Yuhang Ma, Bo Cheng, Shanyuan Liu, Ao Ma 외

Flow-based transformer models for image generation have achieved state-of-the-art performance with larger model parameters, but their inference deployment cost remains high. To enhance inference performance while maintai…

Image Generation

Probabilistic Precipitation Nowcasting with Rectified Flow Transformers

2026-05-29 · Johannes Schusterbauer, Jannik Wiese, Nick Stracke, Timy Phan 외 arxiv

Accurate weather forecasts are essential across various domains and are safety-critical in extreme weather conditions. Compared to simulation-based forecasting, data-driven approaches show greater efficiency, enabling sh…

FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers

2024-12-12 · Yusuf Dalva, Kavana Venkatesh, Pinar Yanardag

Rectified flow models have emerged as a dominant approach in image generation, showcasing impressive capabilities in high-quality image synthesis. However, despite their effectiveness in visual generation, rectified flow…

AttributeDisentanglementImage Generation

FluxSpace: Disentangled Semantic Editing in Rectified Flow Models

2025-01-01 · CVPR 2025 1 · Yusuf Dalva, Kavana Venkatesh, Pinar Yanardag

Rectified flow models have emerged as a dominant approach in image generation, showcasing impressive capabilities in high-quality image synthesis. However, despite their effectiveness in visual generation, rectified …

AttributeDisentanglementImage Generation