paper-with-me

홈 › Papers

SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder

2025-12-12 · Minglei Shi, Haolin Wang, Borui Zhang, Wenzhao Zheng, Bohan Zeng, Ziyang Yuan, Xiaoshi Wu, Yuanxing Zhang, Huan Yang, Xintao Wang, Pengfei Wan, Kun Gai, Jie Zhou, Jiwen Lu arxiv

Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generation. Despite this potential, training large-scale text-to-image diffusion models entirely within the VFM representation space remains largely unexplored. To bridge this gap, we scale the SVG (Self-supervised representations for Visual Generation) framework, proposing SVG-T2I to support high-quality text-to-image synthesis directly in the VFM feature domain. By leveraging a standard text-to-image diffusion pipeline, SVG-T2I achieves competitive performance, reaching 0.75 on GenEval and 85.78 on DPG-Bench. This performance validates the intrinsic representational power of VFMs for generative tasks. We fully open-source the project, including the autoencoder and generation model, together with their training, inference, evaluation pipelines, and pre-trained weights, to facilitate further research in representation-driven visual generation.

📄 PDF Abstract BibTeX arXiv:2512.11749

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LSSGen: Leveraging Latent Space Scaling in Flow and Diffusion for Efficient Text to Image Generation

2025-07-22 · Jyun-Ze Tang, Chih-Fan Hsu, Jeng-Lin Li, Ming-Ching Chang 외 arxiv

Flow matching and diffusion models have shown impressive results in text-to-image generation, producing photorealistic images through an iterative denoising process. A common strategy to speed up synthesis is to perform …

Text-to-Image Generation

Timestep-Aware Diffusion Model for Extreme Image Rescaling

2024-08-17 · Ce Wang, Zhenyu Hu, Wanjie Sun, Zhenzhong Chen

Image rescaling aims to learn the optimal low-resolution (LR) image that can be accurately reconstructed to its original high-resolution (HR) counterpart, providing an efficient image processing and storage method for ul…

DecoderImage Rescalingmodel

DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression

2026-03-13 · Junqi Shi, Ming Lu, Xingchen Li, Anle Ke 외 arxiv

Diffusion-based image compression has recently shown outstanding perceptual fidelity, yet its practicality is hindered by prohibitive sampling overhead and high memory usage. Most existing diffusion codecs employ U-Net a…

Image Compression

Noise Crystallization and Liquid Noise: Zero-shot Video Generation using Image Diffusion Models

2024-10-05 · Muhammad Haaris Khan, Hadrien Reynaud, Bernhard Kainz

Although powerful for image generation, consistent and controllable video is a longstanding problem for diffusion models. Video models require extensive training and computational resources, leading to high costs and lar…

Image GenerationStyle TransferVideo GenerationVideo Style Transfer

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

2026-08-03 · Jiajun Liang, Yucheng Liao, Yukang Cao, Jiazhe Wei 외 hf

Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous …

Text Generation