paper-with-me

홈 › Papers

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

2026-08-17 · Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao, Harry Yang, Steven Hoi arxiv

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.

📄 PDF Abstract BibTeX arXiv:2608.16887

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling

2022-09-04 · CVPR 2023 1 · Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin 외

Masked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) p…

Fill MaskOptical Flow EstimationQuestion AnsweringRetrieval+8

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

2026-08-01 · Marcel Plocher, Bernhard Schölkopf, Andreas Geiger, Gege Gao arxiv

The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumption is particularly questionable for continuous masked generators, wh…

Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation

2026-08-31 · Ran Zhang, Miryam de Lhoneux, Wessel Poelman arxiv

Pixel-based language models (LMs) replace traditional tokenizers by processing rendered images of text, making cross-lingual transfer heavily dependent on the visual and structural properties of writing systems. However,…

Cross-Lingual Transfer

Registers Matter for Pixel-Space Diffusion Transformers

2026-05-15 · Nikita Starodubcev, Ilia Sudakov, Ilya Drobyshevskiy, Artem Babenko 외 arxiv

Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer arch…

Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation

2025-07-03 · François Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe 외

The steep computational cost of diffusion models at inference hinders their use as fast physics emulators. In the context of image and video generation, this computational drawback has been addressed by generating in the…

DiversityVideo Generation