paper-with-me

Papers

ITVTON:Virtual Try-On Diffusion Transformer Model Based on Integrated Image and Text

2025-01-28 · Haifeng Ni

Recent advancements in virtual fitting for characters and clothing have leveraged diffusion models to improve the realism of garment fitting. However, challenges remain in handling complex scenes and poses, which can result in unnatural garment fitting and poorly rendered intricate patterns. In this work, we introduce ITVTON, a novel method that enhances clothing-character interactions by combining clothing and character images along spatial channels as inputs, thereby improving fitting accuracy for the inpainting model. Additionally, we incorporate integrated textual descriptions from multiple images to boost the realism of the generated visual effects. To optimize computational efficiency, we limit training to the attention parameters within a single diffusion transformer (Single-DiT) block. To more rigorously address the complexities of real-world scenarios, we curated training samples from the IGPair dataset, thereby enhancing ITVTON's performance across diverse environments. Extensive experiments demonstrate that ITVTON outperforms baseline methods both qualitatively and quantitatively, setting a new standard for virtual fitting tasks.

📄 PDF Abstract BibTeX arXiv:2501.16757

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

FitVTON: Fit-aware Virtual Try-On via Body-Garment Size Control

2026-06-10 · Yiqun Ning, Ao Shen, Chenhang He, Lei Zhang arxiv

While diffusion-based virtual try-on has achieved impressive visual realism, most methods treat the task as 2D inpainting, prioritizing texture preservation over physical plausibility. Consequently, they often produce pl…

Virtual Try-on

Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On

2024-06-15 · Lingxiao Lu, Shengyi Wu, Haoxuan Sun, Junhong Gou 외

Virtual clothes try-on has emerged as a vital feature in online shopping, offering consumers a critical tool to visualize how clothing fits. In our research, we introduce an innovative approach for virtual clothes try-on…

Virtual Try-on

DiT-VTON: Diffusion Transformer Framework for Unified Multi-Category Virtual Try-On and Virtual Try-All with Integrated Image Editing

2025-10-03 · Qi Li, Shuwen Qiu, Julien Han, Xingzi Xu 외 arxiv

The rapid growth of e-commerce has intensified the demand for Virtual Try-On (VTO) technologies, enabling customers to realistically visualize products overlaid on their own images. Despite recent advances, existing VTO …

Image GenerationVirtual Try-onImage Editing

ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On

2025-03-26 · CVPR 2025 1 · Ji Woo Hong, Tri Ton, Trung X. Pham, Gwanhyeong Koo 외

This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Mask…

DenoisingVirtual Try-on

DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On

2024-12-19 · Wengyi Zhan, Mingbao Lin, Shuicheng Yan, Rongrong Ji

We introduce DiffusionTrend for virtual fashion try-on, which forgoes the need for retraining diffusion models. Using advanced diffusion models, DiffusionTrend harnesses latent information rich in prior information to ca…

DenoisingImage GenerationVirtual Try-on