paper-with-me

Papers

ModaFlow: Modality-Aware Flow Matching for High-Fidelity Virtual Try-On

2026-06-26 · Xiangyu Sai, Meysam Madadi, Sergio Escalera, Yong Xu arxiv

Image-based virtual try-on has emerged as a compelling task in e-commerce and augmented reality, yet existing methods struggle to simultaneously preserve fine garment semantics and adapt to diverse person body geometries under large clothing-body deformations. We present ModaFlow, a modality-aware flow-matching based framework for high-fidelity virtual try-on that achieves precise alignment between textual descriptions and garment appearance. Unlike prior methods that treat multimodal conditions uniformly, ModaFlow introduces a modality-aware guidance scheme: visual garment embeddings extracted by a pretrained image prompt adapter provide deterministic, persistent structural guidance, while textual embeddings generated from garment descriptions are controlled via classifier-free guidance (CFG) with adaptive scaling and zero-initialized velocity. To further enhance flow field accuracy, we propose two regularization losses, cosine similarity and perceptual flow discrimination, that jointly improve directional consistency and perceptual realism of the velocity field. Additionally, a mask manipulation strategy stochastically samples among box, transparent, and relaxed masks during training, simulating diverse occlusion scenarios and enabling robust inference under unpaired settings where only a box mask is available. Experiments show that ModaFlow achieves state-of-the-art results in both qualitative and quantitative evaluations, reducing FID by approximately 30% on paired and 20% on unpaired benchmarks.

📄 PDF Abstract BibTeX arXiv:2606.27773

Code (0)

등록된 구현이 없습니다.

Tasks

Virtual Try-on

Similar Papers 제목 키워드 기반

VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation

2025-08-03 · Xuanran Zhai, Qianyou Zhao, Qiaojun Yu, Ce Hao arxiv

Flow-matching-based policies have recently emerged as a promising approach for learning-based robot manipulation, offering significant acceleration in action sampling compared to diffusion-based policies. However, conven…

Robot Manipulation

AsyncCouple-Flow: Asynchronous Cross-Modal Coupling and Flow Matching for Spatio-Temporal Forecasting

2026-09-15 · Zhixiang Wu, Yining Liu, Bo Zhao, Szu-Yu Chen 외 arxiv

Multi-modal spatio-temporal forecasting (MM-STF) supports weather nowcasting, traffic prediction, and earth-system modeling by combining heterogeneous sources such as physical fields, satellite imagery, and in-situ senso…

Semantic SimilarityWeather ForecastingTraffic Prediction

Bi-modality Images Transfer with a Discrete Process Matching Method

2024-09-06 · Zhe Xiong, Qiaoqiao Ding, Xiaoqun Zhang

Recently, medical image synthesis gains more and more popularity, along with the rapid development of generative models. Medical image synthesis aims to generate an unacquired image modality, often from other observed da…

Data AugmentationDiagnosticImage Generation

CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation

2026-01-30 · Wu Songwei, Jiang Zhiduo, Sun Wandong, Xie Guanghu 외 arxiv

Learning long-horizon robotic manipulation requires jointly achieving expressive behavior modeling, real-time inference, and stable execution, which remains challenging for existing generative policies. Diffusion-based a…

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

2025-07-30 · Weide Liu, Wei Zhou, Jun Liu, Ping Hu 외 arxiv

Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matc…

Medical Image Registration3D ReconstructionImage RetrievalImage Matching