paper-with-me

홈 › Papers

MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention

2025-08-05 · Qi Xie, Yongjia Ma, Donglin Di, Xuehao Gao, Xun Yang arxiv

Achieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal identity coherence. To address these limitations, we propose MoCA, a novel Video Diffusion Model built on a Diffusion Transformer (DiT) backbone, incorporating a Mixture of Cross-Attention mechanism inspired by the Mixture-of-Experts paradigm. Our framework improves inter-frame identity consistency by embedding MoCA layers into each DiT block, where Hierarchical Temporal Pooling captures identity features over varying timescales, and Temporal-Aware Cross-Attention Experts dynamically model spatiotemporal relationships. We further incorporate a Latent Video Perceptual Loss to enhance identity coherence and fine-grained details across video frames. To train this model, we collect CelebIPVid, a dataset of 10,000 high-resolution videos from 1,000 diverse individuals, promoting cross-ethnicity generalization. Extensive experiments on CelebIPVid show that MoCA outperforms existing T2V methods by over 5% across Face similarity.

📄 PDF Abstract BibTeX arXiv:2508.03034

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

2026-06-01 · Yuheng Chen, Teng Hu, Yuji Wang, Qingdong He 외 arxiv

Identity-preserving video generation (IPVG) aims to synthesize high-fidelity videos that follow text prompts while faithfully preserving a reference identity. Despite recent progress, existing IPVG methods still struggle…

Text-to-Video Generation

Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair

2026-08-21 · Jiayi Gao, Changcheng Hua, Jiaqi Tang, Yuxin Peng 외 arxiv

Identity-preserving video generation aims to synthesize videos that follow natural-language instructions while maintaining the visual identity of a given subject. Recent commercial video generation models have achieved s…

Text-to-Video GenerationInstruction Following

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

2025-09-01 · Jiayi Gao, Changcheng Hua, Qingchao Chen, Yuxin Peng 외 arxiv

Identity-preserving text-to-video (IPT2V) generation creates videos faithful to both a reference subject image and a text prompt. While fine-tuning large pretrained video diffusion models on ID-matched data achieves stat…

Text-to-Video GenerationImage Enhancement

InpaintHuman: Reconstructing Occluded Humans with Multi-Scale UV Mapping and Identity-Preserving Diffusion Inpainting

2026-01-05 · Jinlong Fan, Shanshan Zhao, Liang Zheng, Jing Zhang 외 arxiv

Reconstructing complete and animatable 3D human avatars from monocular videos remains challenging, particularly under severe occlusions. While 3D Gaussian Splatting has enabled photorealistic human rendering, existing me…

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

2026-05-06 · Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng, Bing Ma 외 arxiv

Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Despite recent progress, existing methods often suffer from significant…

Text-to-Video Generation