paper-with-me

홈 › Papers

Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses

2025-08-13 · Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi, Jiangning Zhang, Chengming Xu, Weijian Cao, Yabiao Wang, Chengjie Wang, Lizhuang Ma arxiv

Current video generation models struggle with identity preservation under large face poses, primarily facing two challenges: the difficulty in exploring an effective mechanism to integrate identity features into DiT architectures, and the lack of targeted coverage of large face poses in existing open-source video datasets. To address these, we present two key innovations. First, we propose Collaborative Face Experts Fusion (CoFE), which dynamically fuses complementary signals from three specialized experts within the DiT backbone: an identity expert that captures cross-pose invariant features, a semantic expert that encodes high-level visual context, and a detail expert that preserves pixel-level attributes such as skin texture and color gradients. Second, we introduce a data curation pipeline comprising three key components: Face Constraints to ensure diverse large-pose coverage, Identity Consistency to maintain stable identity across frames, and Speech Disambiguation to align textual captions with actual speaking behavior. This pipeline yields LaFID-180K, a large-scale dataset of pose-annotated video clips designed for identity-preserving video generation. Experimental results on several benchmarks demonstrate that our approach significantly outperforms state-of-the-art methods in face similarity, FID, and CLIP semantic alignment. Project page: https://rain152.github.io/CoFE/.

📄 PDF Abstract BibTeX arXiv:2508.09476

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention

2025-08-05 · Qi Xie, Yongjia Ma, Donglin Di, Xuehao Gao 외 arxiv

Achieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal id…

Text-to-Video Generation

Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control

2024-05-27 · Zhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu 외

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving for…

Scene GenerationVideo GenerationVideo Synchronization

Decoupled Video Generation with Chain of Training-free Diffusion Model Experts

2024-08-24 · Wenhao Li, Yichao Cao, Xiu Su, Xi Lin 외

Video generation models hold substantial potential in areas such as filmmaking. However, current video diffusion models need high computational costs and produce suboptimal results due to extreme complexity of video gene…

DenoisingVideo Generation

Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion

2025-01-08 · Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie 외

Creating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotemporal inconsistencies and high computational…

DenoisingDiversityVideo DenoisingVideo Generation

Collaborative Diffusion for Multi-Modal Face Generation and Editing

2023-04-20 · CVPR 2023 1 · Ziqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei Liu

Diffusion models arise as a powerful generative tool recently. Despite the great progress, existing diffusion models mainly focus on uni-modal control, i.e., the diffusion process is driven by only one modality of condit…

DenoisingFace Generation