paper-with-me

홈 › Papers

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

2024-03-13 · CVPR 2025 1 · Enric Corona, Andrei Zanfir, Eduard Gabriel Bazavan, Nikos Kolotouros, Thiemo Alldieck, Cristian Sminchisescu

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion diffusion model, and 2) a novel diffusion-based architecture that augments text-to-image models with both spatial and temporal controls. This supports the generation of high quality video of variable length, easily controllable through high-level representations of human faces and bodies. In contrast to previous work, our method does not require training for each person, does not rely on face detection and cropping, generates the complete image (not just the face or the lips), and considers a broad spectrum of scenarios (e.g. visible torso or diverse subject identities) that are critical to correctly synthesize humans who communicate. We also curate MENTOR, a new and diverse dataset with 3d pose and expression annotations, one order of magnitude larger than previous ones (800,000 identities) and with dynamic gestures, on which we train and ablate our main technical contributions. VLOGGER outperforms state-of-the-art methods in three public benchmarks, considering image quality, identity preservation and temporal consistency while also generating upper-body gestures. We analyze the performance of VLOGGER with respect to multiple diversity metrics, showing that our architectural choices and the use of MENTOR benefit training a fair and unbiased model at scale. Finally we show applications in video editing and personalization.

📄 PDF Abstract BibTeX arXiv:2403.08764

Code (0)

등록된 구현이 없습니다.

Tasks

Face DetectionVideo EditingVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Vlogger: Make Your Dream A Vlog

2024-01-17 · CVPR 2024 1 · Shaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang 외

In this work, we present Vlogger, a generic AI system for generating a minute-level video blog (i.e., vlog) of user descriptions. Different from short videos with a few seconds, vlog often contains a complex storyline wi…

Language ModellingLarge Language ModelVideo Generation

ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education

2026-06-17 · Lorenzo Stacchio, Michele Giordano, Daniele Berardini, Primo Zingaretti 외 arxiv

The advent of Generative Artificial Intelligence (GenAI), and in particular Large Language Models (LLMs), is reshaping educational practice, while intensifying ethical debate about its adoption. To date, the dominant par…

Text-based Animatable 3D Avatars with Morphable Model Alignment

2025-04-22 · Yiqian Wu, Malte Prinzler, Xiaogang Jin, Siyu Tang

The generation of high-quality, animatable 3D head avatars from text has enormous potential in content creation applications such as games, movies, and embodied virtual assistants. Current text-to-3D generation methods t…

3D Generation3DGSText to 3D

Instant 3D Human Avatar Generation using Image Diffusion Models

2024-06-11 · Nikos Kolotouros, Thiemo Alldieck, Enric Corona, Eduard Gabriel Bazavan 외

We present AvatarPopUp, a method for fast, high quality 3D human avatar generation from different input modalities, such as images and text prompts and with control over the generated pose and shape. The common theme is …

3D GenerationImage Generation

A Vlogger-augmented Graph Neural Network Model for Micro-video Recommendation

2024-05-28 · Weijiang Lai, Beihong Jin, Beibei Li, Yiyuan Zheng 외

Existing micro-video recommendation models exploit the interactions between users and micro-videos and/or multi-modal information of micro-videos to predict the next micro-video a user will watch, ignoring the informatio…

Contrastive LearningGraph Neural Network