paper-with-me

Papers

Dancing Avatar: Pose and Text-Guided Human Motion Videos Synthesis with Image Diffusion Model

2023-08-15 · Bosheng Qin, Wentao Ye, Qifan Yu, Siliang Tang, Yueting Zhuang

The rising demand for creating lifelike avatars in the digital realm has led to an increased need for generating high-quality human videos guided by textual descriptions and poses. We propose Dancing Avatar, designed to fabricate human motion videos driven by poses and textual cues. Our approach employs a pretrained T2I diffusion model to generate each video frame in an autoregressive fashion. The crux of innovation lies in our adept utilization of the T2I diffusion model for producing video frames successively while preserving contextual relevance. We surmount the hurdles posed by maintaining human character and clothing consistency across varying poses, along with upholding the background's continuity amidst diverse human movements. To ensure consistent human appearances across the entire video, we devise an intra-frame alignment module. This module assimilates text-guided synthesized human character knowledge into the pretrained T2I diffusion model, synergizing insights from ChatGPT. For preserving background continuity, we put forth a background alignment pipeline, amalgamating insights from segment anything and image inpainting techniques. Furthermore, we propose an inter-frame alignment module that draws inspiration from an auto-regressive pipeline to augment temporal consistency between adjacent frames, where the preceding frame guides the synthesis process of the current frame. Comparisons with state-of-the-art methods demonstrate that Dancing Avatar exhibits the capacity to generate human videos with markedly superior quality, both in terms of human and background fidelity, as well as temporal coherence compared to existing state-of-the-art approaches.

📄 PDF Abstract BibTeX arXiv:2308.07749

Code (0)

등록된 구현이 없습니다.

Tasks

Image Inpainting

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MagicAvatar: Multimodal Avatar Generation and Animation

2023-08-28 · Jianfeng Zhang, Hanshu Yan, Zhongcong Xu, Jiashi Feng 외

This report presents MagicAvatar, a framework for multimodal video generation and animation of human avatars. Unlike most existing methods that generate avatar-centric videos directly from multimodal inputs (e.g., text p…

Video Generation

DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models

2023-04-03 · CVPR 2024 1 · Yukang Cao, Yan-Pei Cao, Kai Han, Ying Shan 외

We present DreamAvatar, a text-and-shape guided framework for generating high-quality 3D human avatars with controllable poses. While encouraging results have been reported by recent methods on text-guided 3D common obje…

NeRF

DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion

2024-09-25 · Yukun Huang, Jianan Wang, Ailing Zeng, Zheng-Jun Zha 외

Leveraging pretrained 2D diffusion models and score distillation sampling (SDS), recent methods have shown promising results for text-to-3D avatar generation. However, generating high-quality 3D avatars capable of expres…

Text to 3D

Dancing Points: Synthesizing Ballroom Dancing with Three-Point Inputs

2026-01-05 · Peizhuo Li, Sebastian Starke, Yuting Ye, Olga Sorkine-Hornung arxiv

Ballroom dancing is a structured yet expressive motion category. Its highly diverse movement and complex interactions between leader and follower dancers make the understanding and synthesis challenging. We demonstrate t…

AvatarStudio: High-fidelity and Animatable 3D Avatar Creation from Text

2023-11-29 · Jianfeng Zhang, Xuanmeng Zhang, Huichao Zhang, Jun Hao Liew 외

We study the problem of creating high-fidelity and animatable 3D avatars from only textual descriptions. Existing text-to-avatar methods are either limited to static avatars which cannot be animated or struggle to genera…

NeRF