paper-with-me

홈 › Papers

FICE: Text-Conditioned Fashion Image Editing With Guided GAN Inversion

2023-01-05 · Martin Pernuš, Clinton Fookes, Vitomir Štruc, Simon Dobrišek

Fashion-image editing represents a challenging computer vision task, where the goal is to incorporate selected apparel into a given input image. Most existing techniques, known as Virtual Try-On methods, deal with this task by first selecting an example image of the desired apparel and then transferring the clothing onto the target person. Conversely, in this paper, we consider editing fashion images with text descriptions. Such an approach has several advantages over example-based virtual try-on techniques, e.g.: (i) it does not require an image of the target fashion item, and (ii) it allows the expression of a wide variety of visual concepts through the use of natural language. Existing image-editing methods that work with language inputs are heavily constrained by their requirement for training sets with rich attribute annotations or they are only able to handle simple text descriptions. We address these constraints by proposing a novel text-conditioned editing model, called FICE (Fashion Image CLIP Editing), capable of handling a wide variety of diverse text descriptions to guide the editing procedure. Specifically with FICE, we augment the common GAN inversion process by including semantic, pose-related, and image-level constraints when generating images. We leverage the capabilities of the CLIP model to enforce the semantics, due to its impressive image-text association capabilities. We furthermore propose a latent-code regularization technique that provides the means to better control the fidelity of the synthesized images. We validate FICE through rigorous experiments on a combination of VITON images and Fashion-Gen text descriptions and in comparison with several state-of-the-art text-conditioned image editing approaches. Experimental results demonstrate FICE generates highly realistic fashion images and leads to stronger editing performance than existing competing approaches.

📄 PDF Abstract BibTeX arXiv:2301.02110

Code (1)

martinpernus/fice 공식 구현 pytorch

Tasks

AttributeVirtual Try-on

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing

2024-03-21 · Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini 외

Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay between clothing and the human body. In the c…

DenoisingVirtual Try-on

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

2023-04-04 · ICCV 2023 1 · Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia 외

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision…

Multimodal fashion image editing

SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

2024-05-01 · Burak Can Biner, Farrin Marouf Sofian, Umur Berkay Karakaş, Duygu Ceylan 외

We are witnessing a revolution in conditional image synthesis with the recent success of large scale text-to-image generation methods. This success also opens up new opportunities in controlling the generation and editin…

Image GenerationText to Image GenerationText-to-Image Generation

MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion

2024-12-28 · Zechao Zhan, Dehong Gao, Jinxia Zhang, Jiale Huang 외

Text-guided image editing model has achieved great success in general domain. However, directly applying these models to the fashion domain may encounter two issues: (1) Inaccurate localization of editing region; (2) Wea…

Large Language Modeltext-guided-image-editing

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

2024-09-02 · Xiaolong Wang, Zhi-Qi Cheng, Jue Wang, Xiaojiang Peng

Fashion image editing is a crucial tool for designers to convey their creative ideas by visualizing design concepts interactively. Current fashion image editing techniques, though advanced with multimodal prompts and pow…

Image GenerationLanguage ModellingLarge Language ModelMultimodal fashion image editing+1