paper-with-me

Papers

Instruct-Video2Avatar: Video-to-Avatar Generation with Instructions

2023-06-05 · Shaoxu Li

We propose a method for synthesizing edited photo-realistic digital avatars with text instructions. Given a short monocular RGB video and text instructions, our method uses an image-conditioned diffusion model to edit one head image and uses the video stylization method to accomplish the editing of other head images. Through iterative training and update (three times or more), our method synthesizes edited photo-realistic animatable 3D neural head avatars with a deformable neural radiance field head synthesis method. In quantitative and qualitative studies on various subjects, our method outperforms state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2306.02903

Code (1)

lsx0101/instruct-video2avatar 공식 구현

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

2024-05-24 · Yuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu 외

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the…

Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

2025-09-11 · Yikang Ding, Jiwen Liu, Wenyuan Zhang, Zekun Wang 외 arxiv

Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-level tracking driven by acoustic or visual…

Domain GeneralizationVideo Generation

MagicAvatar: Multimodal Avatar Generation and Animation

2023-08-28 · Jianfeng Zhang, Hanshu Yan, Zhongcong Xu, Jiashi Feng 외

This report presents MagicAvatar, a framework for multimodal video generation and animation of human avatars. Unlike most existing methods that generate avatar-centric videos directly from multimodal inputs (e.g., text p…

Video Generation

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

2026-09-10 · Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang 외 hf

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video g…

Instruction FollowingVideo Generation

GAIA: Zero-shot Talking Avatar Generation

2023-11-26 · Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang 외

Zero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics such as warping-based motion representat…

Diversity