paper-with-me

Papers

Text-to-Image Generation via Implicit Visual Guidance and Hypernetwork

2022-08-17 · Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, John Collomosse

We develop an approach for text-to-image generation that embraces additional retrieval images, driven by a combination of implicit visual guidance loss and generative objectives. Unlike most existing text-to-image generation methods which merely take the text as input, our method dynamically feeds cross-modal search results into a unified training stage, hence improving the quality, controllability and diversity of generation results. We propose a novel hypernetwork modulated visual-text encoding scheme to predict the weight update of the encoding layer, enabling effective transfer from visual information (e.g. layout, content) into the corresponding latent domain. Experimental results show that our model guided with additional retrieval visual data outperforms existing GAN-based models. On COCO dataset, we achieve better FID of $9.13$ with up to $3.5 \times$ fewer generator parameters, compared with the state-of-the-art method.

📄 PDF Abstract BibTeX arXiv:2208.08493

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage GenerationRetrievalText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

HyperNetwork A HyperNetwork is a network that generates weights for a main network. The behavior of the main network is the same with any usual neural network: it learns to map some raw…

Similar Papers 제목 키워드 기반

The Silent Prompt: Initial Noise as Implicit Guidance for Goal-Driven Image Generation

2024-12-06 · Ruoyu Wang, Huayang Huang, Ye Zhu, Olga Russakovsky 외

Text-to-image synthesis (T2I) has advanced remarkably with the emergence of large-scale diffusion models. In the conventional setup, the text prompt provides explicit, user-defined guidance, directing the generation proc…

DenoisingImage Generation

Implicit and Explicit Language Guidance for Diffusion-based Visual Perception

2024-04-11 · Hefeng Wang, Jiale Cao, Jin Xie, Aiping Yang 외

Text-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high-quality images with rich texture and re…

Depth EstimationImage GenerationSemantic Segmentation

FlexiTex: Enhancing Texture Generation with Visual Guidance

2024-09-19 · Dadong Jiang, Xianghui Yang, Zibo Zhao, Sheng Zhang 외

Recent texture generation methods achieve impressive results due to the powerful generative prior they leverage from large-scale text-to-image diffusion models. However, abstract textual prompts are limited in providing …

Texture Synthesis

Masked Generative Story Transformer with Character Guidance and Caption Augmentation

2024-03-13 · Christos Papadimitriou, Giorgos Filandrianos, Maria Lymperaiou, Giorgos Stamou

Story Visualization (SV) is a challenging generative vision task, that requires both visual quality and consistency between different frames in generated image sequences. Previous approaches either employ some kind of me…

Language ModelingLanguage ModellingLarge Language ModelStory Visualization

Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation

2026-02-03 · Yijia Xu, Zihao Wang, Haokun Gui, Jinshi Cui arxiv

Multi-subject image generation aims to synthesize images that faithfully preserve the identities of multiple reference subjects while following textual instructions. However, existing methods often suffer from identity i…

Image Generation