paper-with-me

Papers

Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

2024-09-04 · Jianwen Jiang, Chao Liang, Jiaqi Yang, Gaojie Lin, Tianyun Zhong, Yanbo Zheng

With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to the limited control of audio signals in driving human motion, existing methods often add auxiliary spatial signals to stabilize movements, which may compromise the naturalness and freedom of motion. In this paper, we propose an end-to-end audio-only conditioned video diffusion model named Loopy. Specifically, we designed an inter- and intra-clip temporal module and an audio-to-latents module, enabling the model to leverage long-term motion information from the data to learn natural motion patterns and improving audio-portrait movement correlation. This method removes the need for manually specified spatial motion templates used in existing methods to constrain motion during inference. Extensive experiments show that Loopy outperforms recent audio-driven portrait diffusion models, delivering more lifelike and high-quality results across various scenarios.

📄 PDF Abstract BibTeX arXiv:2409.02634

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control

2024-12-26 · Wenzhang Sun, Xiang Li, Donglin Di, Zhuding Liang 외

Recently, animating portrait images using audio input is a popular task. Creating lifelike talking head videos requires flexible and natural movements, including facial and head dynamics, camera motion, realistic light a…

DiversityTalking Head Generation

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

2024-01-16 · Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang 외

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simulta…

3D ReconstructionFace GenerationSuper-ResolutionTalking Face Generation

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation

2026-06-29 · Juncheng Ma, Yuxuan Du, Yanan Sun, Zhening Xing 외 arxiv

Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial inference latency. Although training-free diffusion caching accelerates in…

ReliTalk: Relightable Talking Portrait Generation from a Single Video

2023-09-05 · Haonan Qiu, Zhaoxi Chen, Yuming Jiang, Hang Zhou 외

Recent years have witnessed great progress in creating vivid audio-driven portraits from monocular videos. However, how to seamlessly adapt the created video avatars to other scenarios with different backgrounds and ligh…

Single-Image Portrait Relighting

MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices

2024-07-08 · CVPR 2025 1 · Jianwen Jiang, Gaojie Lin, Zhengkun Rong, Chao Liang 외

Existing neural head avatars methods have achieved significant progress in the image quality and motion range of portrait animation. However, these methods neglect the computational overhead, and to the best of our knowl…

Image GenerationPortrait Animation