paper-with-me

Papers

AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding

2024-05-06 · Tao Liu, Feilong Chen, Shuai Fan, Chenpeng Du, Qi Chen, Xie Chen, Kai Yu

The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait. Unlike existing models that primarily focus on verbal cues such as lip synchronization and fail to capture the complex dynamics of facial expressions and nonverbal cues, AniTalker employs a universal motion representation. This innovative representation effectively captures a wide range of facial dynamics, including subtle expressions and head movements. AniTalker enhances motion depiction through two self-supervised learning strategies: the first involves reconstructing target video frames from source frames within the same identity to learn subtle motion representations, and the second develops an identity encoder using metric learning while actively minimizing mutual information between the identity and motion encoders. This approach ensures that the motion representation is dynamic and devoid of identity-specific details, significantly reducing the need for labeled data. Additionally, the integration of a diffusion model with a variance adapter allows for the generation of diverse and controllable facial animations. This method not only demonstrates AniTalker's capability to create detailed and realistic facial movements but also underscores its potential in crafting dynamic avatars for real-world applications. Synthetic results can be viewed at https://github.com/X-LANCE/AniTalker.

📄 PDF Abstract BibTeX arXiv:2405.03121

Code (1)

x-lance/anitalker 공식 구현 pytorch

Tasks

Metric LearningSelf-Supervised Learning

Methods 이 논문이 사용한 방법론

Adapter 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Audio-Driven Talking Face Generation with Diverse yet Realistic Facial Animations

2023-04-18 · Rongliang Wu, Yingchen Yu, Fangneng Zhan, Jiahui Zhang 외

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to th…

Face GenerationTalking Face Generation

LetsTalk: Latent Diffusion Transformer for Talking Video Synthesis

2024-11-24 · Haojie Zhang, Zhihao Liang, Ruibo Fu, Zhengqi Wen 외

Portrait image animation using audio has rapidly advanced, enabling the creation of increasingly realistic and expressive animated faces. The challenges of this multimodality-guided video generation task involve fusing v…

DiversityImage AnimationVideo Generation

VividDream: Generating 3D Scene with Ambient Dynamics

2024-05-30 · Yao-Chih Lee, Yi-Ting Chen, Andrew Wang, Ting-Hsuan Liao 외

We introduce VividDream, a method for generating explorable 4D scenes with ambient dynamics from a single input image or text prompt. VividDream first expands an input image into a static 3D point cloud through iterative…

PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation

2024-12-10 · Fatemeh Nazarieh, ZhenHua Feng, Diptesh Kanojia, Muhammad Awais 외

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects suc…

Face GenerationTalking Face Generation

DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

2023-12-15 · Yifeng Ma, Shiwei Zhang, Jiayu Wang, Xiang Wang 외

Emotional talking head generation has attracted growing attention. Previous methods, which are mainly GAN-based, still struggle to consistently produce satisfactory results across diverse emotions and cannot conveniently…

DenoisingTalking Head Generation