paper-with-me

홈 › Papers

FashionComposer: Compositional Fashion Image Generation

2024-12-18 · Sihui Ji, Yiyang Wang, Xi Chen, Xiaogang Xu, Hao Luo, Hengshuang Zhao

We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and supports personalizing the appearance, pose, and figure of the human and assigning multiple garments in one pass. To achieve this, we first develop a universal framework capable of handling diverse input modalities. We construct scaled training data to enhance the model's robust compositional capabilities. To accommodate multiple reference images (garments and faces) seamlessly, we organize these references in a single image as an "asset library" and employ a reference UNet to extract appearance features. To inject the appearance features into the correct pixels in the generated result, we propose subject-binding attention. It binds the appearance features from different "assets" with the corresponding text features. In this way, the model could understand each asset according to their semantics, supporting arbitrary numbers and types of reference images. As a comprehensive solution, FashionComposer also supports many other applications like human album generation, diverse virtual try-on tasks, etc.

📄 PDF Abstract BibTeX arXiv:2412.14168

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationVirtual Try-on

Similar Papers 제목 키워드 기반

LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing

2025-07-30 · Federico Girella, Davide Talon, Ziyue Liu, Zanxi Ruan 외 arxiv

Fashion design is a complex creative process that blends visual and textual expressions. Designers convey ideas through sketches, which define spatial structure and design elements, and textual descriptions, capturing ma…

Image Generation

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

2026-06-14 · Long-Bao Nguyen, Quang-Khai Tran, Tam V. Nguyen, Minh-Triet Tran 외 arxiv

Multi-subject reference-based image generation requires jointly preserving multiple human identities, binding per-person objects and fashion items, and respecting a specified background scene, a regime where current diff…

Image Generation

Evaluating Attribute Confusion in Fashion Text-to-Image Generation

2025-07-09 · Ziyue Liu, Federico Girella, Yiming Wang, Davide Talon

Despite the rapid advances in Text-to-Image (T2I) generation models, their evaluation remains challenging in domains like fashion, involving complex compositional generation. Recent automated T2I evaluation methods lever…

Attributecross-modal alignmentImage GenerationQuestion Answering+5

Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual Abduction

2025-01-01 · CVPR 2025 1 · Shanshan Huang, Haoxuan Li, Chunyuan Zheng, Mingyuan Ge 외

Fashion image editing is a valuable tool for designers to convey their creative ideas by visualizing design concepts. With the recent advances in text editing methods, significant progress has been made in fashion im…

counterfactualCounterfactual ReasoningDenoising

Text-guided 3D Human Generation from 2D Collections

2023-05-23 · Tsu-Jui Fu, Wenhan Xiong, Yixin Nie, Jingyu Liu 외

3D human modeling has been widely used for engaging interaction in gaming, film, and animation. The customization of these characters is crucial for creativity and scalability, which highlights the importance of controll…

3D geometrytext-to-3d-humanText-to-3D-Human Generation