paper-with-me

홈 › Papers

HeadGAN: One-shot Neural Head Synthesis and Editing

2020-12-15 · ICCV 2021 10 · Michail Christos Doukas, Stefanos Zafeiriou, Viktoriia Sharmanska

Recent attempts to solve the problem of head reenactment using a single reference image have shown promising results. However, most of them either perform poorly in terms of photo-realism, or fail to meet the identity preservation problem, or do not fully transfer the driving pose and expression. We propose HeadGAN, a novel system that conditions synthesis on 3D face representations, which can be extracted from any driving video and adapted to the facial geometry of any reference image, disentangling identity from expression. We further improve mouth movements, by utilising audio features as a complementary input. The 3D face representation enables HeadGAN to be further used as an efficient method for compression and reconstruction and a tool for expression and pose editing.

📄 PDF Abstract BibTeX arXiv:2012.08261

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Free-HeadGAN: Neural Talking Head Synthesis with Explicit Gaze Control

2022-08-03 · Michail Christos Doukas, Evangelos Ververas, Viktoriia Sharmanska, Stefanos Zafeiriou

We present Free-HeadGAN, a person-generic neural talking head synthesis system. We show that modeling faces with sparse 3D facial landmarks are sufficient for achieving state-of-the-art generative performance, without re…

Few-Shot LearningGaze Estimation

BakedAvatar: Baking Neural Fields for Real-Time Head Avatar Synthesis

2023-11-09 · Hao-Bin Duan, Miao Wang, Jin-Chuan Shi, Xu-Chuan Chen 외

Synthesizing photorealistic 4D human head avatars from videos is essential for VR/AR, telepresence, and video game applications. Although existing Neural Radiance Fields (NeRF)-based methods achieve high-fidelity results…

Face ReenactmentNeRF

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

2024-09-11 · Helin Wang, Meng Yu, Jiarui Hai, Chen Chen 외

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot textbased speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer deco…

DecoderSpeech Synthesistext-to-speechText to Speech+1

Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba

2025-10-06 · Baher Mohammad, Magauiya Zhussip, Stamatios Lefkimmiatis arxiv

We introduce MAVE (Mamba with Cross-Attention for Voice Editing and Synthesis), a novel autoregressive architecture for text-conditioned voice editing and high-fidelity text-to-speech (TTS) synthesis, built on a cross-at…

Portrait4D: Learning One-Shot 4D Head Avatar Synthesis using Synthetic Data

2023-11-30 · CVPR 2024 1 · Yu Deng, Duomin Wang, Xiaohang Ren, Xingyu Chen 외

Existing one-shot 4D head synthesis methods usually learn from monocular videos with the aid of 3DMM reconstruction, yet the latter is evenly challenging which restricts them from reasonable 4D head synthesis. We present…

3D Reconstruction