paper-with-me

Papers

VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer

2023-08-09 · Liyang Chen, Zhiyong Wu, Runnan Li, Weihong Bao, Jun Ling, Xu Tan, Sheng Zhao

Current talking face generation methods mainly focus on speech-lip synchronization. However, insufficient investigation on the facial talking style leads to a lifeless and monotonous avatar. Most previous works fail to imitate expressive styles from arbitrary video prompts and ensure the authenticity of the generated video. This paper proposes an unsupervised variational style transfer model (VAST) to vivify the neutral photo-realistic avatars. Our model consists of three key components: a style encoder that extracts facial style representations from the given video prompts; a hybrid facial expression decoder to model accurate speech-related movements; a variational style enhancer that enhances the style space to be highly expressive and meaningful. With our essential designs on facial style learning, our model is able to flexibly capture the expressive facial style from arbitrary video prompts and transfer it onto a personalized image renderer in a zero-shot manner. Experimental results demonstrate the proposed approach contributes to a more vivid talking avatar with higher authenticity and richer expressiveness.

📄 PDF Abstract BibTeX arXiv:2308.04830

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderFace GenerationStyle TransferTalking Face Generation

Methods 이 논문이 사용한 방법론

fail 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

GAIA: Zero-shot Talking Avatar Generation

2023-11-26 · Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang 외

Zero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics such as warping-based motion representat…

Diversity

Generate Your Talking Avatar from Video Reference

2026-04-30 · Zujin Guo, Zhenhui Ye, Yi Ren, Yuanming Li 외 arxiv

Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted, single-view perspective lacks sufficient…

Reinforcement Learning

Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis

2023-06-06 · Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 외

We are interested in a novel task, namely low-resource text-to-talking avatar. Given only a few-minute-long talking person video with the audio track as the training data and arbitrary texts as the driving input, we aim …

Neural Renderingtext-to-speechText to SpeechVideo Generation+1

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

2026-02-02 · Youliang Zhang, Zhengguang Zhou, Zhentao Yu, Ziyao Huang 외 arxiv

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (G…

Video Generation

Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory Guidance

2025-03-28 · Haijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian 외

Pre-trained conditional diffusion models have demonstrated remarkable potential in image editing. However, they often face challenges with temporal consistency, particularly in the talking head domain, where continuous c…