paper-with-me

홈 › Papers

MegActor: Harness the Power of Raw Video for Vivid Portrait Animation

2024-05-31 · Shurong Yang, Huadong Li, Juhao Wu, Minhao Jing, Linze Li, Renhe Ji, Jiajun Liang, Haoqiang Fan

Despite raw driving videos contain richer information on facial expressions than intermediate representations such as landmarks in the field of portrait animation, they are seldom the subject of research. This is due to two challenges inherent in portrait animation driven with raw videos: 1) significant identity leakage; 2) Irrelevant background and facial details such as wrinkles degrade performance. To harnesses the power of the raw videos for vivid portrait animation, we proposed a pioneering conditional diffusion model named as MegActor. First, we introduced a synthetic data generation framework for creating videos with consistent motion and expressions but inconsistent IDs to mitigate the issue of ID leakage. Second, we segmented the foreground and background of the reference image and employed CLIP to encode the background details. This encoded information is then integrated into the network via a text embedding module, thereby ensuring the stability of the background. Finally, we further style transfer the appearance of the reference image to the driving video to eliminate the influence of facial details in the driving videos. Our final model was trained solely on public datasets, achieving results comparable to commercial models. We hope this will help the open-source community.The code is available at https://github.com/megvii-research/MegFaceAnimate.

📄 PDF Abstract BibTeX arXiv:2405.20851

Code (2)

megvii-research/megactor 공식 구현 pytorch
megvii-research/megfaceanimate 공식 구현 pytorch

Tasks

Portrait AnimationStyle TransferSynthetic Data Generation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer

2024-08-27 · Shurong Yang, Huadong Li, Juhao Wu, Minhao Jing 외

Diffusion models have demonstrated superior performance in the field of portrait animation. However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the pot…

Portrait Animation

SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model

2024-09-05 · Weipeng Tan, Chuming Lin, Chengming Xu, Xiaozhong Ji 외

Talking Head Generation (THG), typically driven by audio, is an important and challenging task with broad application prospects in various fields such as digital humans, film production, and virtual reality. While diffus…

DiversityTalking Head Generation

Audio-Driven Emotional Video Portraits

2021-04-15 · CVPR 2021 1 · Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu 외

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important feat…

DisentanglementFace Generation

MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait Animation

2025-01-01 · CVPR 2025 1 · Yukang Lin, Hokit Fung, Jianjin Xu, Zeping Ren 외

Recent portrait animation methods have made significant strides in generating realistic lip synchronization. However, they often lack explicit control over head movements and facial expressions, and cannot produce vi…

Portrait AnimationVideo Generation

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

2026-06-01 · Sicheng Xu, Yu Deng, Shoukang Hu, Yichuan Wang 외 arxiv

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a framework for streamable talking portra…

Video Generation