paper-with-me

홈 › Papers

CanvasGAN: A simple baseline for text to image generation by incrementally patching a canvas

2018-10-05 · Amanpreet Singh, Sharan Agrawal

We propose a new recurrent generative model for generating images from text captions while attending on specific parts of text captions. Our model creates images by incrementally adding patches on a "canvas" while attending on words from text caption at each timestep. Finally, the canvas is passed through an upscaling network to generate images. We also introduce a new method for generating visual-semantic sentence embeddings based on self-attention over text. We compare our model's generated images with those generated Reed et. al.'s model and show that our model is a stronger baseline for text to image generation tasks.

📄 PDF Abstract BibTeX arXiv:1810.02833

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationSentenceSentence EmbeddingsText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control

2022-10-31 · Xiaochuang Han, Sachin Kumar, Yulia Tsvetkov

Despite the growing success of diffusion models in continuous-valued domains (e.g., images), similar efforts for discrete domains such as text have yet to match the performance of autoregressive language models. In this …

DiversityLanguage ModelingLanguage ModellingText Generation

UPainting: Unified Text-to-Image Diffusion Generation with Cross-modal Guidance

2022-10-28 · Wei Li, Xue Xu, Xinyan Xiao, Jiachen Liu 외

Diffusion generative models have recently greatly improved the power of text-conditioned image generation. Existing image generation models mainly include text conditional diffusion model and cross-modal guided diffusion…

Image GenerationImage-text matchingLanguage ModelingLanguage Modelling+1

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

2025-05-29 · Size Wu, Zhonghua Wu, Zerui Gong, Qingyi Tao 외

In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an effi…

Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models

2024-02-21 · Chen Wu, Fernando de la Torre

Text-to-image diffusion models have achieved remarkable performance in image synthesis, while the text interface does not always provide fine-grained control over certain image factors. For instance, changing a single to…

DisentanglementImage GenerationText to Image GenerationText-to-Image Generation

Zero-shot Generation of Coherent Storybook from Plain Text Story using Diffusion Models

2023-02-08 · Hyeonho Jeong, Gihyun Kwon, Jong Chul Ye

Recent advancements in large scale text-to-image models have opened new possibilities for guiding the creation of images through human-devised natural language. However, while prior literature has primarily focused on th…

Language ModelingLanguage ModellingLarge Language Model