paper-with-me

Papers

StyleLipSync: Style-based Personalized Lip-sync Video Generation

2023-04-30 · ICCV 2023 1 · Taekyung Ki, Dongchan Min

In this paper, we present StyleLipSync, a style-based personalized lip-sync video generative model that can generate identity-agnostic lip-synchronizing video from arbitrary audio. To generate a video of arbitrary identities, we leverage expressive lip prior from the semantically rich latent space of a pre-trained StyleGAN, where we can also design a video consistency with a linear transformation. In contrast to the previous lip-sync methods, we introduce pose-aware masking that dynamically locates the mask to improve the naturalness over frames by utilizing a 3D parametric mesh predictor frame by frame. Moreover, we propose a few-shot lip-sync adaptation method for an arbitrary person by introducing a sync regularizer that preserves lip-sync generalization while enhancing the person-specific visual information. Extensive experiments demonstrate that our model can generate accurate lip-sync videos even with the zero-shot setting and enhance characteristics of an unseen face using a few seconds of target video through the proposed adaptation method.

📄 PDF Abstract BibTeX arXiv:2305.00521

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
R1 Regularization R_INLINE_MATH_1 Regularization is a regularization technique and gradient penalty for training [generative adversarial…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Adaptive Instance Normalization 설명 없음
Feedforward Network A Feedforward Network, or a Multilayer Perceptron (MLP), is a neural network with solely densely connected layers. This is the classic neural network architecture of the…
HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
StyleGAN 설명 없음

Similar Papers 제목 키워드 기반

AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data

2024-02-01 · Fu-Yun Wang, Zhaoyang Huang, Weikang Bian, Xiaoyu Shi 외

This paper introduces an effective method for computation-efficient personalized style video generation without requiring access to any personalized video data. It reduces the necessary generation time of similarly sized…

Conditional Image GenerationDenoisingImage GenerationMotion Generation+1

AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation

2023-10-11 · Liyang Chen, Weihong Bao, Shun Lei, Boshi Tang 외

Speech-driven 3D facial animation aims at generating facial movements that are synchronized with the driving speech, which has been widely explored recently. Existing works mostly neglect the person-specific talking styl…

ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer

2024-08-06 · Jiazhi Guan, Zhiliang Xu, Hang Zhou, Kaisiyuan Wang 외

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, th…

Face Swapping

Style-Preserving Lip Sync via Audio-Aware Style Reference

2024-08-10 · Weizhi Zhong, Jichang Li, Yinqi Cai, Ming Li 외

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the uniqu…

PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment

2025-12-27 · Bin Wang, Yang Xu, Huan Zhao, Hao Zhang 외 arxiv

Speech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high lip-synchronization accuracy, existing m…

Talking Head Generation