paper-with-me

홈 › Papers

One-Shot Face Video Re-enactment using Hybrid Latent Spaces of StyleGAN2

2023-02-15 · Trevine Oorloff, Yaser Yacoob

While recent research has progressively overcome the low-resolution constraint of one-shot face video re-enactment with the help of StyleGAN's high-fidelity portrait generation, these approaches rely on at least one of the following: explicit 2D/3D priors, optical flow based warping as motion descriptors, off-the-shelf encoders, etc., which constrain their performance (e.g., inconsistent predictions, inability to capture fine facial details and accessories, poor generalization, artifacts). We propose an end-to-end framework for simultaneously supporting face attribute edits, facial motions and deformations, and facial identity control for video generation. It employs a hybrid latent-space that encodes a given frame into a pair of latents: Identity latent, $\mathcal{W}_{ID}$, and Facial deformation latent, $\mathcal{S}_F$, that respectively reside in the $W+$ and $SS$ spaces of StyleGAN2. Thereby, incorporating the impressive editability-distortion trade-off of $W+$ and the high disentanglement properties of $SS$. These hybrid latents employ the StyleGAN2 generator to achieve high-fidelity face video re-enactment at $1024^2$. Furthermore, the model supports the generation of realistic re-enactment videos with other latent-based semantic edits (e.g., beard, age, make-up, etc.). Qualitative and quantitative analyses performed against state-of-the-art methods demonstrate the superiority of the proposed approach.

📄 PDF Abstract BibTeX arXiv:2302.07848

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDisentanglementOptical Flow EstimationVideo Generation

Methods 이 논문이 사용한 방법론

HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
Path Length Regularization 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
R1 Regularization R_INLINE_MATH_1 Regularization is a regularization technique and gradient penalty for training [generative adversarial…
Weight Demodulation 설명 없음

Similar Papers 제목 키워드 기반

Robust One-Shot Face Video Re-enactment using Hybrid Latent Spaces of StyleGAN2

2023-01-01 · ICCV 2023 1 · Trevine Oorloff, Yaser Yacoob

Recent research on one-shot face re-enactment has progressively overcome the low-resolution constraint with the help of StyleGAN's high-fidelity portrait generation. However, such approaches rely on explicit 2D/3D st…

Attribute

ToonTalker: Cross-Domain Face Reenactment

2023-08-24 · ICCV 2023 1 · Yuan Gong, Yong Zhang, Xiaodong Cun, Fei Yin 외

We target cross-domain face reenactment in this paper, i.e., driving a cartoon image with the video of a real person and vice versa. Recently, many works have focused on one-shot talking face generation to drive a portra…

Face GenerationFace ReenactmentTalking Face Generation

One-shot Face Reenactment

2019-08-05 · Yunxuan Zhang, Siwei Zhang, Yue He, Cheng Li 외

To enable realistic shape (e.g. pose and expression) transfer, existing face reenactment methods rely on a set of target faces for learning subject-specific traits. However, in real-world scenario end-users often only ha…

DecoderFace ReconstructionFace Reenactment

Expressive Talking Head Video Encoding in StyleGAN2 Latent-Space

2022-03-28 · Trevine Oorloff, Yaser Yacoob

While the recent advances in research on video reenactment have yielded promising results, the approaches fall short in capturing the fine, detailed, and expressive facial features (e.g., lip-pressing, mouth puckering, m…

Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model

2025-07-22 · Mingtao Guo, Guanyu Xing, Yanci Zhang, Yanli Liu arxiv

Face reenactment aims to generate realistic talking head videos by transferring motion from a driving video to a static source image while preserving the source identity. Although existing methods based on either implici…