paper-with-me

Papers

ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer

2024-08-06 · Jiazhi Guan, Zhiliang Xu, Hang Zhou, Kaisiyuan Wang, Shengyi He, Zhanwang Zhang, Borong Liang, Haocheng Feng, Errui Ding, Jingtuo Liu, Jingdong Wang, Youjian Zhao, Ziwei Liu

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated models either require long-term videos for clip-specific training or retain visible artifacts. In this paper, we propose a unified and effective framework ReSyncer, that synchronizes generalized audio-visual facial information. The key design is revisiting and rewiring the Style-based generator to efficiently adopt 3D facial dynamics predicted by a principled style-injected Transformer. By simply re-configuring the information insertion mechanisms within the noise and style space, our framework fuses motion and appearance with unified training. Extensive experiments demonstrate that ReSyncer not only produces high-fidelity lip-synced videos according to audio, but also supports multiple appealing properties that are suitable for creating virtual presenters and performers, including fast personalized fine-tuning, video-driven lip-syncing, the transfer of speaking styles, and even face swapping. Resources can be found at https://guanjz20.github.io/projects/ReSyncer.

📄 PDF Abstract BibTeX arXiv:2408.03284

Code (0)

등록된 구현이 없습니다.

Tasks

Face Swapping

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Multi-Head Attention 설명 없음
Attention 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space Learning

2023-05-04 · CVPR 2023 1 · Chao Xu, Junwei Zhu, Jiangning Zhang, Yue Han 외

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical appl…

Face GenerationTalking Face Generation

The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge

2024-10-31 · Dake Guo, Jixun Yao, Xinfa Zhu, Kangxiang Xia 외

This paper presents the NPU-HWC system submitted to the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC). Our system consists of two modules: a speech generator for Track 1 and a backgroun…

Audio GenerationLanguage ModelingLanguage Modelling

StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator

2023-05-09 · CVPR 2023 1 · Jiazhi Guan, Zhanwang Zhang, Hang Zhou, Tianshu Hu 외

Despite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model's generalization ability. Previous studies either require long-term data f…

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

2026-08-31 · Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang 외 hf

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video genera…

Reinforcement LearningVideo Generation

Style Transfer for 2D Talking Head Animation

2023-03-17 · Trong-Thang Pham, Nhat Le, Tuong Do, Hung Nguyen 외

Audio-driven talking head animation is a challenging research topic with many real-world applications. Recent works have focused on creating photo-realistic 2D animation, while learning different talking or singing style…

Style Transfer