paper-with-me

홈 › Papers

REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers

2025-04-14 · Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie, Liang Zheng

In this paper we tackle a fundamental question: "Can we train latent diffusion models together with the variational auto-encoder (VAE) tokenizer in an end-to-end manner?" Traditional deep-learning wisdom dictates that end-to-end training is often preferable when possible. However, for latent diffusion transformers, it is observed that end-to-end training both VAE and diffusion-model using standard diffusion-loss is ineffective, even causing a degradation in final performance. We show that while diffusion loss is ineffective, end-to-end training can be unlocked through the representation-alignment (REPA) loss -- allowing both VAE and diffusion model to be jointly tuned during the training process. Despite its simplicity, the proposed training recipe (REPA-E) shows remarkable performance; speeding up diffusion model training by over 17x and 45x over REPA and vanilla training recipes, respectively. Interestingly, we observe that end-to-end tuning with REPA-E also improves the VAE itself; leading to improved latent space structure and downstream generation performance. In terms of final performance, our approach sets a new state-of-the-art; achieving FID of 1.26 and 1.83 with and without classifier-free guidance on ImageNet 256 x 256. Code is available at https://end2end-diffusion.github.io.

📄 PDF Abstract BibTeX arXiv:2504.10483

Code (2)

End2End-Diffusion/REPA-E pytorch
westlake-repl/leanvae pytorch

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers

2025-04-15 · Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing 외

In this paper we tackle a fundamental question: "Can we train latent diffusion models together with the variational auto-encoder (VAE) tokenizer in an end-to-end manner?" Traditional deep-learning wisdom dictates that en…

Image Generation

Representation Alignment for Just Image Transformers is not Easier than You Think

2026-03-15 · Jaeyo Shin, Jiwook Kim, Hyunjung Shim arxiv

Representation Alignment (REPA) has emerged as a simple way to accelerate Diffusion Transformers training in latent space. At the same time, pixel-space diffusion transformers such as Just image Transformers (JiT) have a…

RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model

2025-12-12 · Guanfang Dong, Luke Schultz, Negar Hassanpour, Chao Gao arxiv

Semantic-rich features from Vision Foundation Models (VFMs) have been leveraged to enhance Latent Diffusion Models (LDMs). However, raw VFM features are typically high-dimensional and redundant, increasing the difficulty…

TextLDM: Language Modeling with Continuous Latent Diffusion

2026-05-08 · Jiaxiu Jiang, Jingjing Ren, Wenbo Li, Bo Wang 외 arxiv

Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesi…

multimodal generationText Generation

Rethinking Cross-Layer Information Routing in Diffusion Transformers

2026-05-20 · Chao Xu, Maohua Li, Qirui Li, Yixuan Xu 외 arxiv

Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has …