paper-with-me

홈 › Papers

MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes

2024-10-09 · Zhenhui Ye, Tianyun Zhong, Yi Ren, Ziyue Jiang, Jiawei Huang, Rongjie Huang, Jinglin Liu, Jinzheng He, Chen Zhang, Zehan Wang, Xize Chen, Xiang Yin, Zhou Zhao

Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Source code and video samples are available at https://mimictalk.github.io .

📄 PDF Abstract BibTeX arXiv:2410.06734

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationNeRFTalking Face Generation

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer

2023-08-09 · Liyang Chen, Zhiyong Wu, Runnan Li, Weihong Bao 외

Current talking face generation methods mainly focus on speech-lip synchronization. However, insufficient investigation on the facial talking style leads to a lifeless and monotonous avatar. Most previous works fail to i…

DecoderFace GenerationStyle TransferTalking Face Generation

AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation

2024-02-25 · Yasheng Sun, Wenqing Chu, Hang Zhou, Kaisiyuan Wang 외

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's sp…

Face GenerationHallucinationTalking Face Generation

One-Shot Pose-Driving Face Animation Platform

2024-07-12 · He Feng, Donglin Di, Yongjia Ma, Wei Chen 외

The objective of face animation is to generate dynamic and expressive talking head videos from a single reference face, utilizing driving conditions derived from either video or audio inputs. Current approaches often req…

Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose

2020-02-24 · Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao 외

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this proble…

3D Face AnimationVideo Generation

MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control

2026-01-30 · Renjie Lu, Xulong Zhang, Xiaoyang Qu, Jianzong Wang 외 arxiv

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic…