paper-with-me

Papers

Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing

2024-03-21 · Alberto Baldrati, Davide Morelli, Marcella Cornia, Marco Bertini, Rita Cucchiara

Fashion illustration is a crucial medium for designers to convey their creative vision and transform design concepts into tangible representations that showcase the interplay between clothing and the human body. In the context of fashion design, computer vision techniques have the potential to enhance and streamline the design process. Departing from prior research primarily focused on virtual try-on, this paper tackles the task of multimodal-conditioned fashion image editing. Our approach aims to generate human-centric fashion images guided by multimodal prompts, including text, human body poses, garment sketches, and fabric textures. To address this problem, we propose extending latent diffusion models to incorporate these multiple modalities and modifying the structure of the denoising network, taking multimodal prompts as input. To condition the proposed architecture on fabric textures, we employ textual inversion techniques and let diverse cross-attention layers of the denoising network attend to textual and texture information, thus incorporating different granularity conditioning details. Given the lack of datasets for the task, we extend two existing fashion datasets, Dress Code and VITON-HD, with multimodal annotations. Experimental evaluations demonstrate the effectiveness of our proposed approach in terms of realism and coherence concerning the provided multimodal inputs.

📄 PDF Abstract BibTeX arXiv:2403.14828

Code (2)

aimagelab/ti-mgd 공식 구현
aimagelab/multimodal-garment-designer pytorch

Tasks

DenoisingVirtual Try-on

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing

2023-04-04 · ICCV 2023 1 · Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia 외

Fashion illustration is used by designers to communicate their vision and to bring the design idea from conceptualization to realization, showing how clothes interact with the human body. In this context, computer vision…

Multimodal fashion image editing

FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion

2024-04-26 · Abhishek Kumar Singh, Ioannis Patras

The rapid evolution of the fashion industry increasingly intersects with technological advancements, particularly through the integration of generative AI. This study introduces a novel generative pipeline designed to tr…

Virtual Try-on

DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing

2024-09-02 · Xiaolong Wang, Zhi-Qi Cheng, Jue Wang, Xiaojiang Peng

Fashion image editing is a crucial tool for designers to convey their creative ideas by visualizing design concepts interactively. Current fashion image editing techniques, though advanced with multimodal prompts and pow…

Image GenerationLanguage ModellingLarge Language ModelMultimodal fashion image editing+1

FashionEngine: Interactive 3D Human Generation and Editing via Multimodal Controls

2024-04-02 · Tao Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu

We present FashionEngine, an interactive 3D human generation and editing system that creates 3D digital humans via user-friendly multimodal controls such as natural languages, visual perceptions, and hand-drawing sketche…

Virtual Try-on

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis

2025-03-19 · CVPR 2025 1 · Jiaqi Liu, Jichao Zahng, Paolo Rota, Nicu Sebe

The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, …

Image Generation