paper-with-me

Papers

GAIA: Zero-shot Talking Avatar Generation

2023-11-26 · Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang, Jialiang Zhu, Kaikai An, Leyi Li, Xu Tan, Chunyu Wang, Han Hu, HsiangTao Wu, Sheng Zhao, Jiang Bian

Zero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics such as warping-based motion representation and 3D Morphable Models, which limit the naturalness and diversity of the generated avatars. In this work, we introduce GAIA (Generative AI for Avatar), which eliminates the domain priors in talking avatar generation. In light of the observation that the speech only drives the motion of the avatar while the appearance of the avatar and the background typically remain the same throughout the entire video, we divide our approach into two stages: 1) disentangling each frame into motion and appearance representations; 2) generating motion sequences conditioned on the speech and reference portrait image. We collect a large-scale high-quality talking avatar dataset and train the model on it with different scales (up to 2B parameters). Experimental results verify the superiority, scalability, and flexibility of GAIA as 1) the resulting model beats previous baseline models in terms of naturalness, diversity, lip-sync quality, and visual quality; 2) the framework is scalable since larger models yield better results; 3) it is general and enables different applications like controllable talking avatar generation and text-instructed avatar generation.

📄 PDF Abstract BibTeX arXiv:2311.15230

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer

2023-08-09 · Liyang Chen, Zhiyong Wu, Runnan Li, Weihong Bao 외

Current talking face generation methods mainly focus on speech-lip synchronization. However, insufficient investigation on the facial talking style leads to a lifeless and monotonous avatar. Most previous works fail to i…

DecoderFace GenerationStyle TransferTalking Face Generation

Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis

2023-06-06 · Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 외

We are interested in a novel task, namely low-resource text-to-talking avatar. Given only a few-minute-long talking person video with the audio track as the training data and arbitrary texts as the driving input, we aim …

Neural Renderingtext-to-speechText to SpeechVideo Generation+1

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

2024-01-16 · Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang 외

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simulta…

3D ReconstructionFace GenerationSuper-ResolutionTalking Face Generation

Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking Avatar

2022-10-13 · Aolan Sun, xulong Zhang, Tiandong Ling, Jianzong Wang 외

Since the beginning of the COVID-19 pandemic, remote conferencing and school-teaching have become important tools. The previous applications aim to save the commuting cost with real-time interactions. However, our applic…

text-to-speechText to Speech

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

2026-02-02 · Youliang Zhang, Zhengguang Zhou, Zhentao Yu, Ziyao Huang 외 arxiv

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (G…

Video Generation