paper-with-me

Papers

Generate Your Talking Avatar from Video Reference

2026-04-30 · Zujin Guo, Zhenhui Ye, Yi Ren, Yuanming Li, Ce Chen, Zhibin Hong, Chen Change Loy arxiv

Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted, single-view perspective lacks sufficient temporal and expression cues, limiting the ability to synthesize high-fidelity talking avatars in customized backgrounds. To this end, we introduce Talking Avatar generation from Video Reference (TAVR), a novel framework that shifts the paradigm by leveraging cross-scene video inputs. To effectively process these extended temporal contexts and bridge cross-scene domain gaps, TAVR integrates a token selection module alongside a comprehensive three-stage training scheme. Specifically, same-scene video pretraining establishes foundational appearance copying, which is subsequently expanded by cross-scene reference fine-tuning for robust cross-scene adaptation. Finally, task-specific reinforcement learning aligns the generated outputs with identity-based rewards to maximize identity similarity. To systematically evaluate cross-scene robustness, we construct a new benchmark comprising 158 carefully curated cross-scene video pairs. Extensive experiments show that TAVR benefits from flexible inference-time video referencing and consistently surpasses existing baselines both quantitatively and qualitatively. This work has been deployed to production. For more related research, please visit \href{https://www.heygen.com/research}{HeyGen Research} and \href{https://www.heygen.com/research/avatar-v-model}{HeyGen Avatar-V}.

📄 PDF Abstract BibTeX arXiv:2604.27918

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style Transfer

2023-08-09 · Liyang Chen, Zhiyong Wu, Runnan Li, Weihong Bao 외

Current talking face generation methods mainly focus on speech-lip synchronization. However, insufficient investigation on the facial talking style leads to a lifeless and monotonous avatar. Most previous works fail to i…

DecoderFace GenerationStyle TransferTalking Face Generation

GAIA: Zero-shot Talking Avatar Generation

2023-11-26 · Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang 외

Zero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics such as warping-based motion representat…

Diversity

Avatar V: Scaling Video-Reference Avatar Video Generation

2026-06-11 · Benjamin Liang, Ce Chen, Desmond Lin, Ivan Somov 외 arxiv

Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies, and expression dynamics, remains an op…

Video GenerationStyle Transfer

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

2024-01-16 · Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang 외

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simulta…

3D ReconstructionFace GenerationSuper-ResolutionTalking Face Generation

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

2026-02-02 · Youliang Zhang, Zhengguang Zhou, Zhentao Yu, Ziyao Huang 외 arxiv

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (G…

Video Generation