paper-with-me

Papers

Avatar V: Scaling Video-Reference Avatar Video Generation

2026-06-11 · Benjamin Liang, Ce Chen, Desmond Lin, Ivan Somov, Jiajun Zhao, Jiewei Yuan, Jingfeng Zhang, Junhao Huang, Nik Nolte, Pedram Haqiqi, Penghan Wang, Rong Yan, Rui Zhang, Sam Prokopchuk, Sivan Wang, Viktor Goriachko, Yi Ren, Yuanming Li, Yutao Chen, Zhenhui Ye, Zhibin Hong, Zilong Nie, Zujin Guo arxiv

Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies, and expression dynamics, remains an open challenge. Existing methods predominantly condition on single static images, which provide insufficient identity information and cannot capture dynamic motion traits, while standard pixel-level objectives underserve the perceptually critical facial regions that determine avatar fidelity. We present Avatar V, a production-scale framework that addresses these limitations through video-reference-conditioned identity modeling. Rather than compressing identity into fixed-size embeddings, the model conditions directly on the full token sequence of a reference video, learning to reproduce both static identity attributes (facial geometry, skin texture) and dynamic behavioral patterns (talking rhythm, micro-expressions) through attention over the reference context. We introduce Sparse Reference Attention, an asymmetric mechanism achieving linear-complexity conditioning on arbitrarily long references; a motion representation stream enabling closed-loop talking style transfer; and an identity-aware super-resolution refiner inheriting the full reference conditioning. These are supported by a data engine curating 100M+ training clips from 50M raw videos, and a five-stage training pipeline with flow matching pre-training, personality fine-tuning, two-phase distillation (>10x acceleration), and RLHF alignment, deployed across thousands of GPUs. Avatar V generates 1080p videos of unlimited duration, achieving state-of-the-art identity preservation, lip synchronization, and generation quality on our cross-scene benchmark, consistently outperforming leading systems including Seedance 2.0, Kling O3 Pro, Veo 3.1, and OmniHuman 1.5 in both automated metrics and human evaluation.

📄 PDF Abstract BibTeX arXiv:2606.13872

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationStyle Transfer

Similar Papers 제목 키워드 기반

Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual Reality

2024-08-13 · Yu-Chih Chen, Avinab Saha, ALEXANDRE CHAPIRO, Christian Häne 외

We study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected …

Video CompressionVideo Quality AssessmentVisual Question Answering (VQA)

Generate Your Talking Avatar from Video Reference

2026-04-30 · Zujin Guo, Zhenhui Ye, Yi Ren, Yuanming Li 외 arxiv

Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted, single-view perspective lacks sufficient…

Reinforcement Learning

MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars

2025-10-14 · Felix Taubner, Ruihang Zhang, Mathieu Tuli, Sherwin Bahmani 외 arxiv

Digital human avatars aim to simulate the dynamic appearance of humans in virtual environments, enabling immersive experiences across gaming, film, virtual reality, and more. However, the conventional process for creatin…

Video Generation

OPHAvatars: One-shot Photo-realistic Head Avatars

2023-07-18 · Shaoxu Li

We propose a method for synthesizing photo-realistic digital avatars from only one portrait as the reference. Given a portrait, our method synthesizes a coarse talking head video using driving keypoints features. And wit…

Blind Face Restoration

LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar Animation

2026-03-02 · Hualiang Wei, Shunran Jia, Jialun Liu, Wenhui Li arxiv

We present LiftAvatar, a new paradigm that completes sparse monocular observations in kinematic space (e.g., facial expressions and head pose) and uses the completed signals to drive high-fidelity avatar animation. LiftA…