paper-with-me

Papers

V-Warper: Appearance-Consistent Video Diffusion Personalization via Value Warping

2025-12-13 · Hyunkoo Lee, Wooseok Jang, Jini Yang, Taehwan Kim, Sangoh Kim, Sangwon Jung, Seungryong Kim arxiv

Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-based finetuning or large-scale video datasets, which impose substantial computational cost and are difficult to scale. Furthermore, they still struggle to maintain fine-grained appearance consistency across frames. To address these limitations, we introduce V-Warper, a training-free coarse-to-fine personalization framework for transformer-based video diffusion models. The framework enhances fine-grained identity fidelity without requiring any additional video training. (1) A lightweight coarse appearance adaptation stage leverages only a small set of reference images, which are already required for the task. This step encodes global subject identity through image-only LoRA and subject-embedding adaptation. (2) A inference-time fine appearance injection stage refines visual fidelity by computing semantic correspondences from RoPE-free mid-layer query--key features. These correspondences guide the warping of appearance-rich value representations into semantically aligned regions of the generation process, with masking ensuring spatial reliability. V-Warper significantly improves appearance fidelity while preserving prompt alignment and motion dynamics, and it achieves these gains efficiently without large-scale video finetuning.

📄 PDF Abstract BibTeX arXiv:2512.12375

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

2025-01-02 · Yuanpeng Tu, Hao Luo, Xi Chen, Sihui Ji 외

Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of the reference object and accurately model…

Talking Head GenerationVideo GenerationVirtual Try-on

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

2025-07-08 · Zhenghao Zhang, Junchao Liao, Xiangyu Meng, Long Qin 외

Recent advances in diffusion transformer models for motion-guided video generation, such as Tora, have shown significant progress. In this paper, we present Tora2, an enhanced version of Tora, which introduces several de…

Video Generation

ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation

2025-12-08 · Ziyang Mai, Yu-Wing Tai arxiv

Text-to-video (T2V) generation has advanced rapidly, yet maintaining consistent character identities across scenes remains a major challenge. Existing personalization methods often focus on facial identity but fail to pr…

Text-to-Video Generation

ChronoLock: Protecting Videos from Unauthorized Text-to-Video Personalization

2026-06-19 · Jiaming He, Jiashu Zhang, Guanyu Hou, Shuhan Ye 외 arxiv

Text-to-video (T2V) diffusion models have made it increasingly easy to synthesize realistic and temporally coherent videos, while recent personalization techniques allow such models to imitate a specific subject, style, …

Learning Temporal Pose Estimation from Sparsely-Labeled Videos

2019-06-06 · NeurIPS 2019 12 · Gedas Bertasius, Christoph Feichtenhofer, Du Tran, Jianbo Shi 외

Modern approaches for multi-person pose estimation in video require large amounts of dense annotations. However, labeling every frame in a video is costly and labor intensive. To reduce the need for dense annotations, we…

Multi-Person Pose EstimationOptical Flow EstimationPose Estimation