paper-with-me

홈 › Papers

SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training

2026-01-25 · Mengmeng Wang, Dengyang Jiang, Liuzhuozheng Li, Yucheng Lin, Guojiang Shen, Xiangjie Kong, Yong Liu, Guang Dai, Jingdong Wang arxiv

Denoising-based diffusion transformers, despite their strong generation performance, suffer from inefficient training convergence. Existing methods addressing this issue, such as REPA (relying on external representation encoders) or SRA (requiring dual-model setups), inevitably incur heavy computational overhead during training due to external dependencies. To tackle these challenges, this paper proposes SRA 2, a lightweight intrinsic guidance framework for efficient diffusion training. SRA 2 leverages off-the-shelf pre-trained Variational Autoencoder (VAE) features: their reconstruction property ensures inherent encoding of visual priors like rich texture details, structural patterns, and basic semantic information. Specifically, SRA 2 aligns the intermediate latent features of diffusion transformers with VAE features via a lightweight projection layer, supervised by a feature alignment loss. This design accelerates training without extra representation encoders or dual-model maintenance, resulting in a simple yet effective pipeline. Extensive experiments demonstrate that SRA 2 improves both generation quality and training convergence speed compared to vanilla diffusion transformers, matches or outperforms state-of-the-art acceleration methods, and incurs merely 4% extra GFLOPs with zero additional cost for external guidance models.

📄 PDF Abstract BibTeX arXiv:2601.17830

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Text-to-speech synthesis based on latent variable conversion using diffusion probabilistic model and variational autoencoder

2022-12-16 · Yusuke Yasuda, Tomoki Toda

Text-to-speech synthesis (TTS) is a task to convert texts into speech. Two of the factors that have been driving TTS are the advancements of probabilistic models and latent representation learning. We propose a TTS metho…

Representation LearningSpeech Synthesistext-to-speechText to Speech+1

Latent Diffusion Model without Variational Autoencoder

2025-10-17 · Minglei Shi, Haolin Wang, Wenzhao Zheng, Ziyang Yuan 외 arxiv

Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with variational autoencoders (VAEs). While effective for high-fidelity synthesis, this VAE+diffusion paradigm suffers fr…

USP: Unified Self-Supervised Pretraining for Image Generation and Understanding

2025-03-08 · Xiangxiang Chu, Renda Li, Yong Wang

Recent studies have highlighted the interplay between diffusion models and representation learning. Intermediate representations from diffusion models can be leveraged for downstream visual tasks, while self-supervised v…

Image GenerationRepresentation Learning

SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder

2025-12-12 · Minglei Shi, Haolin Wang, Borui Zhang, Wenzhao Zheng 외 arxiv

Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generation. Despite this potential, training lar…

Diffusion Variational Autoencoders

2019-01-25 · Luis A. Pérez Rey, Vlado Menkovski, Jacobus W. Portegies

A standard Variational Autoencoder, with a Euclidean latent space, is structurally incapable of capturing topological properties of certain datasets. To remove topological obstructions, we introduce Diffusion Variational…