paper-with-me

홈 › Papers

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

2025-07-29 · He Feng, Yongjia Ma, Donglin Di, Lei Fan, Tonghua Su, Xiangqian Wu arxiv

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking styles. Existing diffusion-based portrait animation methods primarily focus on lip synchronization or static emotion transformation, often overlooking dynamic styles such as head movements. Moreover, most of these methods rely on a dual U-Net architecture, which preserves identity consistency but incurs additional computational overhead. To this end, we propose DiTalker, a unified DiT-based framework for speaking style-controllable portrait animation. We design a Style-Emotion Encoding Module that employs two separate branches: a style branch extracting identity-specific style information (e.g., head poses and movements), and an emotion branch extracting identity-agnostic emotion features. We further introduce an Audio-Style Fusion Module that decouples audio and speaking styles via two parallel cross-attention layers, using these features to guide the animation process. To enhance the quality of results, we adopt and modify two optimization constraints: one to improve lip synchronization and the other to preserve fine-grained identity and background details. Extensive experiments demonstrate the superiority of DiTalker in terms of lip synchronization and speaking style controllability. Project Page: https://thenameishope.github.io/DiTalker/

📄 PDF Abstract BibTeX arXiv:2508.06511

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer

2024-08-03 · Yihong Lin, Zhaoxin Fan, Xianjia Wu, Lingyu Xiong 외

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have ach…

DiversityTalking Head Generation

MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation

2024-03-28 · Seyeon Kim, Siyoon Jin, JiHye Park, Kihong Kim 외

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. Howe…

Talking Head Generation

EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents

2026-04-19 · Yu Zhang, Kaiyuan Shen, Yang Li arxiv

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking b…

StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads

2024-09-14 · Suzhen Wang, Yifeng Ma, Yu Ding, Zhipeng Hu 외

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail t…

Face GenerationTalking Face Generation

Learning Singing From Speech

2019-12-20 · Liqiang Zhang, Chengzhu Yu, Heng Lu, Chao Weng 외

We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unifi…

Speech SynthesisVoice Conversion