paper-with-me

홈 › Papers

Laminating Representation Autoencoders for Efficient Diffusion

2026-02-04 · Ramón Calvo-González, François Fleuret arxiv

Recent work has shown that diffusion models can generate high-quality images by operating directly on SSL patch features rather than pixel-space latents. However, the dense patch grids from encoders like DINOv2 contain significant redundancy, making diffusion needlessly expensive. We introduce FlatDINO, a variational autoencoder that compresses this representation into a one-dimensional sequence of just 32 continuous tokens -an 8x reduction in sequence length and 48x compression in total dimensionality. On ImageNet 256x256, a DiT-XL trained on FlatDINO latents achieves a gFID of 1.80 with classifier-free guidance while requiring 8x fewer FLOPs per forward pass and up to 4.5x fewer FLOPs per training step compared to diffusion on uncompressed DINOv2 features. These are preliminary results and this work is in progress.

📄 PDF Abstract BibTeX arXiv:2602.04873

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diffusion Models as Masked Autoencoders

2023-04-06 · ICCV 2023 1 · Chen Wei, Karttikeya Mangalam, Po-Yao Huang, Yanghao Li 외

There has been a longstanding belief that generation can facilitate a true understanding of visual data. In line with this, we revisit generatively pre-training visual representations in light of recent interest in denoi…

DenoisingImage Inpainting

Denoising Diffusion Autoencoders are Unified Self-supervised Learners

2023-03-17 · ICCV 2023 1 · Weilai Xiang, Hongyu Yang, Di Huang, Yunhong Wang

Inspired by recent advances in diffusion models, which are reminiscent of denoising autoencoders, we investigate whether they can acquire discriminative representations for classification via generative pre-training. Thi…

Contrastive LearningDenoisingImage GenerationLinear evaluation+2

Hierarchical Diffusion Autoencoders and Disentangled Image Manipulation

2023-04-24 · Zeyu Lu, Chengyue Wu, Xinyuan Chen, Yaohui Wang 외

Diffusion models have attained impressive visual quality for image synthesis. However, how to interpret and manipulate the latent space of diffusion models has not been extensively explored. Prior work diffusion autoenco…

Image GenerationImage ManipulationImage Reconstruction

On Designing Diffusion Autoencoders for Efficient Generation and Representation Learning

2025-05-30 · Magdalena Proszewska, Nikolay Malkin, N. Siddharth

Diffusion autoencoders (DAs) are variants of diffusion generative models that use an input-dependent latent variable to capture representations alongside the diffusion process. These representations, to varying extents, …

DenoisingRepresentation Learning

Slice estimation in diffusion MRI of neonatal and fetal brains in image and spherical harmonics domains using autoencoders

2022-08-29 · Hamza Kebiri, Gabriel Girard, Yasser Aleman-Gomez, Thomas Yu 외

Diffusion MRI (dMRI) of the developing brain can provide valuable insights into the white matter development. However, slice thickness in fetal dMRI is typically high (i.e., 3-5 mm) to freeze the in-plane motion, which r…

AnatomyDiffusion MRI