paper-with-me

홈 › Papers

Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism

2024-12-13 · Jun Zheng, Jing Wang, Fuwei Zhao, Xujie Zhang, Xiaodan Liang

Video try-on stands as a promising area for its tremendous real-world potential. Previous research on video try-on has primarily focused on transferring product clothing images to videos with simple human poses, while performing poorly with complex movements. To better preserve clothing details, those approaches are armed with an additional garment encoder, resulting in higher computational resource consumption. The primary challenges in this domain are twofold: (1) leveraging the garment encoder's capabilities in video try-on while lowering computational requirements; (2) ensuring temporal consistency in the synthesis of human body parts, especially during rapid movements. To tackle these issues, we propose a novel video try-on framework based on Diffusion Transformer(DiT), named Dynamic Try-On. To reduce computational overhead, we adopt a straightforward approach by utilizing the DiT backbone itself as the garment encoder and employing a dynamic feature fusion module to store and integrate garment features. To ensure temporal consistency of human body parts, we introduce a limb-aware dynamic attention module that enforces the DiT backbone to focus on the regions of human limbs during the denoising process. Extensive experiments demonstrate the superiority of Dynamic Try-On in generating stable and smooth try-on results, even for videos featuring complicated human postures.

📄 PDF Abstract BibTeX arXiv:2412.09822

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingVirtual Try-on

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
ADOPT Please enter a description about the method here
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Focus 설명 없음

Similar Papers 제목 키워드 기반

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

2025-01-20 · Zheng Chong, Wenqing Zhang, Shiyue Zhang, Jun Zheng 외

Virtual try-on (VTON) technology has gained attention due to its potential to transform online retail by enabling realistic clothing visualization of images and videos. However, most existing methods struggle to achieve …

Video GenerationVirtual Try-on

Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

2025-05-22 · CVPR 2025 1 · Dong Li, Wenqi Zhong, Wei Yu, Yingwei Pan 외

Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and phys…

DenoisingVirtual Try-on

Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation

2025-12-25 · Steven Xiao, Xindi Zhang, Dechao Meng, Qi Wang 외 arxiv

Real-time portrait animation is essential for interactive applications such as virtual assistants and live avatars, requiring high visual fidelity, temporal coherence, ultra-low latency, and responsive control from dynam…

Video Generation

RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency

2025-01-15 · Siqi Li, Zhengkai Jiang, Jiawei Zhou, Zhihong Liu 외

Virtual try-on has emerged as a pivotal task at the intersection of computer vision and fashion, aimed at digitally simulating how clothing items fit on the human body. Despite notable progress in single-image virtual tr…

Virtual Try-on

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation

2025-10-10 · Chuanrui Zhang, Zhengxian Wu, Guanxing Lu, Yansong Tang 외 arxiv

Learned world models hold significant potential as neural simulators for robotic manipulation. However, prevalent 2D video-based models inherently lack the spatial and kinematic reasoning crucial for physical interaction…

Reinforcement Learning