paper-with-me

Papers

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

2024-12-05 · Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan, Zhenxiong Tan, Jiahao Lu, Chuanxin Tang, Bo An, Shuicheng Yan

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing natural, audio-aligned expressions in generated talking videos remain significant challenges. To address these challenges, we propose Memory-guided EMOtion-aware diffusion (MEMO), an end-to-end audio-driven portrait animation approach to generate identity-consistent and expressive talking videos. Our approach is built around two key modules: (1) a memory-guided temporal module, which enhances long-term identity consistency and motion smoothness by developing memory states to store information from a longer past context to guide temporal modeling via linear attention; and (2) an emotion-aware audio module, which replaces traditional cross attention with multi-modal attention to enhance audio-video interaction, while detecting emotions from audio to refine facial expressions via emotion adaptive layer norm. Extensive quantitative and qualitative results demonstrate that MEMO generates more realistic talking videos across diverse image and audio types, outperforming state-of-the-art methods in overall quality, audio-lip synchronization, identity consistency, and expression-emotion alignment.

📄 PDF Abstract BibTeX arXiv:2412.04448

Code (0)

등록된 구현이 없습니다.

Tasks

Portrait AnimationVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

TalkingPose: Efficient Face and Gesture Animation with Feedback-guided Diffusion Model

2025-11-30 · Alireza Javanmardi, Pragati Jaiswal, Tewodros Amberbir Habtegebrial, Christen Millerdurai 외 arxiv

Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of …

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

2024-11-23 · CVPR 2025 1 · Haotian Wang, Yuzhe Weng, Yueyan Li, Zilu Guo 외

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk …

Talking Head Generation

Versatile Multimodal Controls for Expressive Talking Human Animation

2025-03-10 · Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li 외

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not onl…

Human Animation

Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion

2025-02-11 · Xingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang 외

Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches …

AttributeDisentanglementFace GenerationPortrait Animation+1

AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation

2024-02-25 · Yasheng Sun, Wenqing Chu, Hang Zhou, Kaisiyuan Wang 외

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's sp…

Face GenerationHallucinationTalking Face Generation