paper-with-me

홈 › Papers

ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search

2025-11-10 · Zhenjie Liu, Jianzhang Lu, Renjie Lu, Cong Liang, Shangfei Wang arxiv

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily stem from entangled appearance-motion representations and unstable inference strategies. In this paper, we introduce \textbf{ConsistTalk}, a novel intensity-controllable and temporally consistent talking head generation framework with diffusion noise search inference. First, we propose \textbf{an optical flow-guided temporal module (OFT)} that decouples motion features from static appearance by leveraging facial optical flow, thereby reducing visual flicker and improving temporal consistency. Second, we present an \textbf{Audio-to-Intensity (A2I) model} obtained through multimodal teacher-student knowledge distillation. By transforming audio and facial velocity features into a frame-wise intensity sequence, the A2I model enables joint modeling of audio and visual motion, resulting in more natural dynamics. This further enables fine-grained, frame-wise control of motion dynamics while maintaining tight audio-visual synchronization. Third, we introduce a \textbf{diffusion noise initialization strategy (IC-Init)}. By enforcing explicit constraints on background coherence and motion continuity during inference-time noise search, we achieve better identity preservation and refine motion dynamics compared to the current autoregressive strategy. Extensive experiments demonstrate that ConsistTalk significantly outperforms prior methods in reducing flicker, preserving identity, and delivering temporally stable, high-fidelity talking head videos.

📄 PDF Abstract BibTeX arXiv:2511.06833

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head GenerationKnowledge Distillation

Similar Papers 제목 키워드 기반

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

2026-07-01 · Haijie Yang, Zhenyu Zhang, Yixuan Dong, Jianjun Qian 외 arxiv

Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially unde…

DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis

2025-10-12 · Peiyin Chen, Zhuowei Yang, Hui Feng, Sheng Jiang 외 arxiv

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-mat…

Dynamic Neural Textures: Generating Talking-Face Videos with Continuously Controllable Expressions

2022-04-13 · Zipeng Ye, Zhiyao Sun, Yu-Hui Wen, Yanan sun 외

Recently, talking-face video generation has received considerable attention. So far most methods generate results with neutral expressions or expressions that are implicitly determined by neural networks in an uncontroll…

Video Generation

Talking-head Generation with Rhythmic Head Motion

2020-07-16 · Lele Chen, Guofeng Cui, Celong Liu, Zhong Li 외

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably suc…

Talking Head Generation

FreeTalk: Emotional Topology-Free 3D Talking Heads

2026-03-16 · Federico Nocentini, Thomas Besnier, Claudio Ferrari, Stefano Berretti 외 arxiv

Speech-driven 3D facial animation has advanced rapidly, yet most approaches remain tied to registered template meshes, preventing effective deployment on raw 3D scans with arbitrary topology. At the same time, modeling c…