Talking Head Generation
7개 벤치마크 · 논문 159편 · 이 태스크의 논문 보기 →
Benchmarks
VoxCeleb2 - 1-shot learning
VoxCeleb1 - 1-shot learning
VoxCeleb1 - 32-shot learning
VoxCeleb1 - 8-shot learning
VoxCeleb2 - 8-shot learning
Most implemented
Few-Shot Adversarial Learning of Realistic Neural Talking Head Models
A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
MakeItTalk: Speaker-Aware Talking-Head Animation
SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation
Papers
Decoupled Self-Forcing Distillation for Streaming Talking Head Generation
Streaming talking-head generation produces each frame as its driving audio arrives, yet fidelity and efficiency have so far pulled in opposite directions: end-to-end methods condition a video diffusion model on audio dir…
Talking Head GenerationLeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer …
Talking Head GenerationVideo GenerationTemporally-Aligned Evaluation for Audio-Driven Talking Head Generation
Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temporal correspondence between generated and reference videos. This assum…
Talking Head GenerationCapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation
Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predef…
Talking Head GenerationMoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation
Talking-head generation requires joint modeling of identity, head pose, facial expression, and mouth dynamics. Existing methods typically address only a subset of these factors, and rely on fixed-weight or heuristic fusi…
Talking Head GenerationAsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
Diffusion-based talking head generation has achieved remarkable visual quality, yet scaling it to long-term videos remains challenging. The widely adopted chunk-wise paradigm introduces two fundamental failures: (1) temp…
Talking Head GenerationKnowledge Distillation