paper-with-me

홈 › Papers

Identity-Consistent Video Generation under Large Facial-Angle Variations

2026-03-22 · Bin Hu, Zipeng Qi, Guoxi Huang, Zunnan Xu, Ruicheng Zhang, Chongjie Ye, Jun Zhou, Xiu Li, Jingdong Wang arxiv

Single-view reference-to-video methods often struggle to preserve identity consistency under large facial-angle variations. This limitation naturally motivates the incorporation of multi-view facial references. However, simply introducing additional reference images exacerbates the \textit{copy-paste} problem, particularly the \textbf{\textit{view-dependent copy-paste}} artifact, which reduces facial motion naturalness. Although cross-paired data can alleviate this issue, collecting such data is costly. To balance the consistency and naturalness, we propose $\mathrm{Mv}^2\mathrm{ID}$, a multi-view conditioned framework under in-paired supervision. We introduce a region-masking training strategy to prevent shortcut learning and extract essential identity features by encouraging the model to aggregate complementary identity cues across views. In addition, we design a reference decoupled-RoPE mechanism that assigns distinct positional encoding to video and conditioning tokens for better modeling of their heterogeneous properties. Furthermore, we construct a large-scale dataset with diverse facial-angle variations and propose dedicated evaluation metrics for identity consistency and motion naturalness. Extensive experiments demonstrate that our method significantly improves identity consistency while maintaining motion naturalness, outperforming existing approaches trained with cross-paired data.

📄 PDF Abstract BibTeX arXiv:2603.21299

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

2026-06-09 · Cong Wang, Zhentao Yu, Hongmei Wang, Weicong Liang 외 arxiv

Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-view reference input offers a natural solution, progress remains constr…

Spatial ReasoningVideo Generation

WildActor: Unconstrained Identity-Preserving Video Generation

2026-02-28 · Qin Guo, Tianyu Yang, Xuanhua He, Fei Shen 외 arxiv

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. …

Video Generation

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

2026-05-06 · Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng, Bing Ma 외 arxiv

Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Despite recent progress, existing methods often suffer from significant…

Text-to-Video Generation

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

2025-10-16 · Xiangyu Meng, Zixian Zhang, Zhenghao Zhang, Junchao Liao 외 arxiv

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent ident…

Reinforcement LearningVideo Generation

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

2026-02-12 · Maomao Li, Zhen Li, Kaipeng Zhang, Guosheng Yin 외 arxiv

Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, t…

Contrastive LearningVideo Generation