paper-with-me

홈 › Papers

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

2026-07-10 · Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park arxiv

Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g., tucked in or untucked, open or closed), and spatial placement on the body. We address this gap with two complementary contributions. First, we define and solve Visual-Instance-Prompt Segmentation via VIP-SAM: given a flatlay image of a garment, segment that specific instance in a photograph of a person wearing it. This is an instance-level task, distinct from the typically studied category-level segmentation. Second, we introduce CtrlVTON, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body. VIP-SAM and CtrlVTON each achieve state-of-the-art results on their respective tasks. In particular, CtrlVTON generates images that follow user-provided layouts far more faithfully than the strongest proprietary editing systems while matching them on garment fidelity.

📄 PDF Abstract BibTeX arXiv:2607.09362

Code (0)

등록된 구현이 없습니다.

Tasks

Virtual Try-onImage Editing

Similar Papers 제목 키워드 기반

InstanceControl: Controllable Complex Image Generation without Instance Labeling

2026-06-30 · Xiaoyu Liu, Huan Wang, Fan Li, Zhixin Wang 외 hf

Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with co…

Image Generation

ConsistCompose: Unified Multimodal Layout Control for Image Composition

2025-11-23 · Xuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 외 arxiv

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counter…

Visual GroundingImage Generation

Layout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint

2024-10-20 · Junwei Zhou, Xueting Li, Lu Qi, Ming-Hsuan Yang

We present Layout-Your-3D, a framework that allows controllable and compositional 3D generation from text prompts. Existing text-to-3D methods often struggle to generate assets with plausible object interactions or requi…

3D GenerationText to 3D

Contrastive Learning with Prompt-derived Virtual Semantic Prototypes for Unsupervised Sentence Embedding

2022-11-07 · Jiali Zeng, Yongjing Yin, Yufan Jiang, Shuangzhi Wu 외

Contrastive learning has become a new paradigm for unsupervised sentence embeddings. Previous studies focus on instance-wise contrastive learning, attempting to construct positive pairs with textual data augmentation. In…

ClusteringContrastive LearningData AugmentationSemantic Textual Similarity+4

Brain on the 3D Visual Art through Virtual Reality; Introducing Neuro-Art in a Case Investigation

2019-04-14

The reciprocal impact of applied neuroscience and cognitive studies on humanities has been extensive and growing over the past 30 years of research. Studies on neuroaesthetics have provided novel insights in visual arts,…