paper-with-me

홈 › Papers

Towards Open-World Text-Guided Face Image Generation and Manipulation

2021-04-18 · Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu

The existing text-guided image synthesis methods can only produce limited quality results with at most \mbox{$\text{256}^2$} resolution and the textual instructions are constrained in a small Corpus. In this work, we propose a unified framework for both face image generation and manipulation that produces diverse and high-quality images with an unprecedented resolution at 1024 from multimodal inputs. More importantly, our method supports open-world scenarios, including both image and text, without any re-training, fine-tuning, or post-processing. To be specific, we propose a brand new paradigm of text-guided image generation and manipulation based on the superior characteristics of a pretrained GAN model. Our proposed paradigm includes two novel strategies. The first strategy is to train a text encoder to obtain latent codes that align with the hierarchically semantic of the aforementioned pretrained GAN model. The second strategy is to directly optimize the latent codes in the latent space of the pretrained GAN model with guidance from a pretrained language model. The latent codes can be randomly sampled from a prior distribution or inverted from a given image, which provides inherent supports for both image generation and manipulation from multi-modal inputs, such as sketches or semantic labels, with textual guidance. To facilitate text-guided multi-modal synthesis, we propose the Multi-Modal CelebA-HQ, a large-scale dataset consisting of real face images and corresponding semantic segmentation map, sketch, and textual descriptions. Extensive experiments on the introduced dataset demonstrate the superior performance of our proposed method. Code and data are available at https://github.com/weihaox/TediGAN.

📄 PDF Abstract BibTeX arXiv:2104.08910

Code (2)

weihaox/TediGAN 공식 구현 pytorch
IIGROUP/TediGAN pytorch

Tasks

Image GenerationLanguage ModellingSemantic SegmentationText-to-Image Generation

Similar Papers 제목 키워드 기반

Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images

2023-08-31 · ICCV 2023 1 · Cuican Yu, Guansong Lu, Yihan Zeng, Jian Sun 외

Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie, and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape…

3D Shape GenerationContrastive Learningcross-modal alignmentFace Generation+2

WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark

2026-02-06 · Wang Lin, Feng Wang, Majun Zhang, Wentao Hu 외 arxiv

Recent advances in image editing models have demonstrated remarkable capabilities in executing explicit instructions, such as attribute manipulation, style transfer, and pose synthesis. However, these models often face c…

Instruction FollowingStyle TransferImage Editing

Pluralistic Aging Diffusion Autoencoder

2023-03-20 · ICCV 2023 1 · Peipei Li, Rui Wang, Huaibo Huang, Ran He 외

Face aging is an ill-posed problem because multiple plausible aging patterns may correspond to a given input. Most existing methods often produce one deterministic estimation. This paper proposes a novel CLIP-driven Plur…

DenoisingDiversity

ChatFace: Chat-Guided Real Face Editing via Diffusion Latent Space Manipulation

2023-05-24 · Dongxu Yue, Qin Guo, Munan Ning, Jiaxi Cui 외

Editing real facial images is a crucial task in computer vision with significant demand in various real-world applications. While GAN-based methods have showed potential in manipulating images especially when combined wi…

AttributeImage Reconstruction

LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization

2025-03-15 · Yuchen Deng, Haibin Ling, Bingyao Huang

We propose LAPIG, a language guided projector image generation method with surface adaptation and stylization. LAPIG consists of a projector-camera system and a target textured projection surface. LAPIG takes the user te…

Image Generation