paper-with-me

Papers

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

2026-08-17 · Kwan Yun, Serin Yoon, Sunjin Jung, Jung Eun Yoo, Inyup Lee, Junyong Noh arxiv

We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (CsF) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing "no motion"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, AnyTalk_{RT}, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at https://serin-yoon.github.io/projects/anytalk/.

📄 PDF Abstract BibTeX arXiv:2608.16143

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models

2025-03-11 · CVPR 2025 1 · Kwan Yun, Seokhyeon Hong, Chaelin Kim, Junyong Noh

Despite recent advancements in learning-based motion in-betweening, a key limitation has been overlooked: the requirement for character-specific datasets. In this work, we introduce AnyMoLe, a novel method that addresses…

Motion Generationmotion in-betweeningMotion Synthesis

Speech Driven Tongue Animation

2022-01-01 · CVPR 2022 1 · Salvador Medina, Denis Tome, Carsten Stoll, Mark Tiede 외

Advances in speech driven animation techniques allow the creation of convincing animations for virtual characters solely from audio data. Many existing approaches focus on facial and lip motion and they often do not …

Decoder

MoCha: Towards Movie-Grade Talking Character Synthesis

2025-03-30 · Cong Wei, Bo Sun, Haoyu Ma, Ji Hou 외

Recent advancements in video generation have achieved impressive motion realism, yet they often overlook character-driven storytelling, a crucial task for automated film, animation generation. We introduce Talking Charac…

Video Generation

AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement

2025-11-28 · Zhizhou Zhong, Yicheng Ji, Zhe Kong, Yiying Liu 외 arxiv

Recently, multi-person video generation has started to gain prominence. While a few preliminary works have explored audio-driven multi-person talking video generation, they often face challenges due to the high costs of …

Video Generation

DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context Learning

2026-01-29 · Mingshuang Luo, Shuang Liang, Zhengkun Rong, Yuxuan Luo 외 arxiv

Character image animation aims to synthesize high-fidelity videos by transferring motion from a driving sequence to a static reference image. Despite recent advancements, existing methods suffer from two fundamental chal…

Domain Generalization