paper-with-me

Papers

Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance

2024-10-17 · Jiwan Hur, Dong-Jae Lee, Gyojin Han, Jaehyun Choi, Yunho Jeon, Junmo Kim

Masked generative models (MGMs) have shown impressive generative ability while providing an order of magnitude efficient sampling steps compared to continuous diffusion models. However, MGMs still underperform in image synthesis compared to recent well-developed continuous diffusion models with similar size in terms of quality and diversity of generated samples. A key factor in the performance of continuous diffusion models stems from the guidance methods, which enhance the sample quality at the expense of diversity. In this paper, we extend these guidance methods to generalized guidance formulation for MGMs and propose a self-guidance sampling method, which leads to better generation quality. The proposed approach leverages an auxiliary task for semantic smoothing in vector-quantized token space, analogous to the Gaussian blur in continuous pixel space. Equipped with the parameter-efficient fine-tuning method and high-temperature sampling, MGMs with the proposed self-guidance achieve a superior quality-diversity trade-off, outperforming existing sampling methods in MGMs with more efficient training and sampling costs. Extensive experiments with the various sampling hyperparameters confirm the effectiveness of the proposed self-guidance.

📄 PDF Abstract BibTeX arXiv:2410.13136

Code (1)

jiwanhur/unlockmgm 공식 구현 jax

Tasks

DiversityImage Generationparameter-efficient fine-tuning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis

2022-11-16 · CVPR 2023 1 · Tianhong Li, Huiwen Chang, Shlok Kumar Mishra, Han Zhang 외

Generative modeling and representation learning are two key tasks in computer vision. However, these models are typically trained independently, which ignores the potential for each task to help the other, and leads to t…

Image GenerationRepresentation LearningUnconditional Image Generation

Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis

2024-10-10 · Jinbin Bai, Tian Ye, Wei Chow, Enxin Song 외

We present Meissonic, which elevates non-autoregressive masked image modeling (MIM) text-to-image to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architec…

Feature CompressionImage Generation

StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis

2023-01-23 · Axel Sauer, Tero Karras, Samuli Laine, Andreas Geiger 외

Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model families such as diffusion and autoregressive m…

Image GenerationText-to-Image Generation

Science-T2I: Addressing Scientific Illusions in Image Synthesis

2025-04-17 · CVPR 2025 1 · Jialuo Li, Wenhao Chai, Xingyu Fu, Haiyang Xu 외

We present a novel approach to integrating scientific knowledge into generative models, enhancing their realism and consistency in image synthesis. First, we introduce Science-T2I, an expert-annotated adversarial dataset…

Image Generation

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

2026-06-29 · Shufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu 외 arxiv

We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion…

Image Generation