paper-with-me

홈 › Papers

Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off

2025-08-06 · Seungyong Lee, Jeong-gi Kwak arxiv

Virtual try-on aims to synthesize a realistic image of a person wearing a target garment, but accurately modeling garment-body correspondence remains a persistent challenge, especially under pose and appearance variation. In this paper, we propose Voost - a unified and scalable framework that jointly learns virtual try-on and try-off with a single diffusion transformer. By modeling both tasks jointly, Voost enables each garment-person pair to supervise both directions and supports flexible conditioning over generation direction and garment category, enhancing garment-body relational reasoning without task-specific networks, auxiliary losses, or additional labels. In addition, we introduce two inference-time techniques: attention temperature scaling for robustness to resolution or mask variation, and self-corrective sampling that leverages bidirectional consistency between tasks. Extensive experiments demonstrate that Voost achieves state-of-the-art results on both try-on and try-off benchmarks, consistently outperforming strong baselines in alignment accuracy, visual fidelity, and generalization.

📄 PDF Abstract BibTeX arXiv:2508.04825

Code (0)

등록된 구현이 없습니다.

Tasks

Relational ReasoningVirtual Try-on

Similar Papers 제목 키워드 기반

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

2026-05-05 · Lin Song, Wenbo Li, Guoqing Ma, Wei Tang 외 arxiv

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language M…

Text-to-Image GenerationImage Editing

Text-to-3D Generation with Bidirectional Diffusion using both 2D and 3D priors

2023-12-07 · CVPR 2024 1 · Lihe Ding, Shaocong Dong, Zhanpeng Huang, Zibin Wang 외

Most 3D generation research focuses on up-projecting 2D foundation models into the 3D space, either by minimizing 2D Score Distillation Sampling (SDS) loss or fine-tuning on multi-view datasets. Without explicit 3D prior…

3D GenerationDiversityText to 3DTexture Synthesis

DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

2024-10-10 · Jiatao Gu, Yuyang Wang, Yizhe Zhang, Qihang Zhang 외

Diffusion models have become the dominant approach for visual generation. They are trained by denoising a Markovian process which gradually adds noise to the input. We argue that the Markovian property limits the model's…

DenoisingImage GenerationQuantizationText to Image Generation+1

Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation

2024-05-24 · Shentong Mo, Yapeng Tian

In recent developments, the Mamba architecture, known for its selective state space approach, has shown potential in the efficient modeling of long sequences. However, its application in image generation remains underexp…

Image GenerationMambaVideo Generation

D3LM: A Discrete DNA Diffusion Language Model for Bidirectional DNA Understanding and Generation

2026-03-02 · Zhao Yang, Hengchang Liu, Chuan Cao, Bing Su arxiv

Early DNA foundation models adopted BERT-style training, achieving good performance on DNA understanding tasks but lacking generative capabilities. Recent autoregressive models enable DNA generation, but employ left-to-r…

Representation Learning