paper-with-me

Papers

GenSync: A Generalized Talking Head Framework for Audio-driven Multi-Subject Lip-Sync using 3D Gaussian Splatting

2025-05-03 · Anushka Agarwal, Muhammad Yusuf Hassan, Talha Chafekar

We introduce GenSync, a novel framework for multi-identity lip-synced video synthesis using 3D Gaussian Splatting. Unlike most existing 3D methods that require training a new model for each identity , GenSync learns a unified network that synthesizes lip-synced videos for multiple speakers. By incorporating a Disentanglement Module, our approach separates identity-specific features from audio representations, enabling efficient multi-identity video synthesis. This design reduces computational overhead and achieves 6.8x faster training compared to state-of-the-art models, while maintaining high lip-sync accuracy and visual quality.

📄 PDF Abstract BibTeX arXiv:2505.01928

Code (0)

등록된 구현이 없습니다.

Tasks

Disentanglement

Similar Papers 제목 키워드 기반

DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation

2023-01-10 · CVPR 2023 1 · Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 외

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. Ho…

DenoisingTalking Head Generation

GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

2023-01-31 · Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 외

Generating photo-realistic video portrait with arbitrary speech audio is a crucial problem in film-making and virtual reality. Recently, several works explore the usage of neural radiance field in this task to improve 3D…

Face GenerationLip ReadingNeRFTalking Face Generation+1

SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

2025-01-24 · Yujian Liu, Shidang Xu, Jing Guo, Dingbin Wang 외

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for appli…

NeRF

Style Transfer for 2D Talking Head Animation

2023-03-17 · Trong-Thang Pham, Nhat Le, Tuong Do, Hung Nguyen 외

Audio-driven talking head animation is a challenging research topic with many real-world applications. Recent works have focused on creating photo-realistic 2D animation, while learning different talking or singing style…

Style Transfer

DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis

2024-09-16 · Fa-Ting Hong, Yunfei Liu, Yu Li, Changyin Zhou 외

Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. H…

Talking Head Generation