CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.
Code (0)
등록된 구현이 없습니다.
Tasks
Virtual Try-onImage EditingSimilar Papers 제목 키워드 기반
InstanceControl: Controllable Complex Image Generation without Instance Labeling
Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with co…
Image GenerationConsistCompose: Unified Multimodal Layout Control for Image Composition
Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counter…
Visual GroundingImage GenerationLayout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint
We present Layout-Your-3D, a framework that allows controllable and compositional 3D generation from text prompts. Existing text-to-3D methods often struggle to generate assets with plausible object interactions or requi…
3D GenerationText to 3DContrastive Learning with Prompt-derived Virtual Semantic Prototypes for Unsupervised Sentence Embedding
Contrastive learning has become a new paradigm for unsupervised sentence embeddings. Previous studies focus on instance-wise contrastive learning, attempting to construct positive pairs with textual data augmentation. In…
ClusteringContrastive LearningData AugmentationSemantic Textual Similarity+4Brain on the 3D Visual Art through Virtual Reality; Introducing Neuro-Art in a Case Investigation
The reciprocal impact of applied neuroscience and cognitive studies on humanities has been extensive and growing over the past 30 years of research. Studies on neuroaesthetics have provided novel insights in visual arts,…