paper-with-me

홈 › Papers

Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding

2026-01-04 · Yixuan Lai, He Wang, Kun Zhou, Tianjia Shao arxiv

Producing prompt-faithful videos that preserve a user-specified identity remains challenging: models need to extrapolate facial dynamics from sparse reference while balancing the tension between identity preservation and motion naturalness. Conditioning on a single image completely ignores the temporal signature, which leads to pose-locked motions, unnatural warping, and "average" faces when viewpoints and expressions change. To this end, we introduce an identity-conditioned variant of a diffusion-transformer video generator which uses a short reference video rather than a single portrait. Our key idea is to incorporate the dynamics in the reference. A short clip reveals subject-specific patterns, e.g., how smiles form, across poses and lighting. From this clip, a Sinkhorn-routed encoder learns compact identity tokens that capture characteristic dynamics while remaining pretrained backbone-compatible. Despite adding only lightweight conditioning, the approach consistently improves identity retention under large pose changes and expressive facial behavior, while maintaining prompt faithfulness and visual realism across diverse subjects and prompts.

📄 PDF Abstract BibTeX arXiv:2601.01352

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

2026-06-01 · Yuheng Chen, Teng Hu, Yuji Wang, Qingdong He 외 arxiv

Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference identity. Despite recent progress, existing IPVG methods still struggle…

Text-to-Video Generation

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

2025-09-01 · Jiayi Gao, Changcheng Hua, Qingchao Chen, Yuxin Peng 외 arxiv

Identity-preserving text-to-video (IPT2V) generation creates videos faithful to both a reference subject image and a text prompt. While fine-tuning large pretrained video diffusion models on ID-matched data achieves stat…

Text-to-Video GenerationImage Enhancement

Vera: Identity-Faithful Human Subject-to-Video Generation

2026-07-22 · Yulong Xu, Xinyue Liu, Shujuan Li, huafeng shi 외 arxiv

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may a…

Video Generation

KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation

2026-08-17 · Jianjie Luo, Yiming Zhong, Haoming Shen, Yupeng Xiao 외 arxiv

Identity-preserving video generation (IPVG) requires synthesizing videos that are faithful to both reference subjects and text prompts. Existing methods are often hindered by high tuning costs or limited input-level enha…

Video Generation

AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References

2026-03-26 · Jiahao Wang, Hualian Sheng, Sijia Cai, Yuxiao Yang 외 arxiv

Identity-preserving video generation offers powerful tools for creative expression, allowing users to customize videos featuring their beloved characters. However, prevailing methods are typically designed and optimized …

Reinforcement LearningVideo Generation