paper-with-me

Papers

Text-Conditioned Sampling Framework for Text-to-Image Generation with Masked Generative Models

2023-04-04 · ICCV 2023 1 · Jaewoong Lee, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon, Yunji Kim, Jin-Hwa Kim, Jung-Woo Ha, Sung Ju Hwang

Token-based masked generative models are gaining popularity for their fast inference time with parallel decoding. While recent token-based approaches achieve competitive performance to diffusion-based models, their generation performance is still suboptimal as they sample multiple tokens simultaneously without considering the dependence among them. We empirically investigate this problem and propose a learnable sampling model, Text-Conditioned Token Selection (TCTS), to select optimal tokens via localized supervision with text information. TCTS improves not only the image quality but also the semantic alignment of the generated images with the given texts. To further improve the image quality, we introduce a cohesive sampling strategy, Frequency Adaptive Sampling (FAS), to each group of tokens divided according to the self-attention maps. We validate the efficacy of TCTS combined with FAS with various generative tasks, demonstrating that it significantly outperforms the baselines in image-text alignment and image quality. Our text-conditioned sampling framework further reduces the original inference time by more than 50% without modifying the original generative model.

📄 PDF Abstract BibTeX arXiv:2304.01515

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Simultaneous Image-to-Zero and Zero-to-Noise: Diffusion Models with Analytical Image Attenuation

2023-06-23 · Yuhang Huang, Zheng Qin, Xinwang Liu, Kai Xu

Recent studies have demonstrated that the forward diffusion process is crucial for the effectiveness of diffusion models in terms of generative quality and sampling efficiency. We propose incorporating an analytical imag…

DenoisingEdge DetectionImage GenerationImage Inpainting+2

SceneGenie: Scene Graph Guided Diffusion Models for Image Synthesis

2023-04-28 · Azade Farshad, Yousef Yeganeh, Yu Chi, Chengzhi Shen 외

Text-conditioned image generation has made significant progress in recent years with generative adversarial networks and more recently, diffusion models. While diffusion models conditioned on text prompts have produced i…

Image GenerationImage Generation from Scene GraphsSegmentationText to Image Generation+1

There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

2026-08-28 · Gabe Guo, Elon Litman, Thanawat Sornwanee, Jose Blanchet 외 arxiv

Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some …

UIFace: Unleashing Inherent Model Capabilities to Enhance Intra-Class Diversity in Synthetic Face Recognition

2025-02-27 · Xiao Lin, Yuge Huang, Jianqing Xu, Yuxi Mi 외

Face recognition (FR) stands as one of the most crucial applications in computer vision. The accuracy of FR models has significantly improved in recent years due to the availability of large-scale human face datasets. Ho…

DiversityFace RecognitionSynthetic Face Recognition

clip2latent: Text driven sampling of a pre-trained StyleGAN using denoising diffusion and CLIP

2022-10-05 · Justin N. M. Pinkney, Chuan Li

We introduce a new method to efficiently create text-to-image models from a pre-trained CLIP and StyleGAN. It enables text driven sampling with an existing generative model without any external data or fine-tuning. This …

Denoising