paper-with-me

홈 › Papers

UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation

2026-03-02 · Hebeizi Li, Zihao Liang, Benyuan Sun, Zihao Yin, Xiao Sha, Chenliang Wang, Yi Yang arxiv

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and performance, we introduce UniTalking, a unified, end-to-end diffusion framework for generating high-fidelity speech and lip-synchronized video. At its core, our framework employs Multi-Modal Transformer Blocks to explicitly model the fine-grained temporal correspondence between audio and video latent tokens via a shared self-attention mechanism. By leveraging powerful priors from a pre-trained video generation model, our framework ensures state-of-the-art visual fidelity while enabling efficient training. Furthermore, UniTalking incorporates a personalized voice cloning capability, allowing the generation of speech in a target style from a brief audio reference. Qualitative and quantitative results demonstrate that our method produces highly realistic talking portraits, achieving superior performance over existing open-source approaches in lip-sync accuracy, audio naturalness, and overall perceptual quality.

📄 PDF Abstract BibTeX arXiv:2603.01418

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads

2024-09-14 · Suzhen Wang, Yifeng Ma, Yu Ding, Zhipeng Hu 외

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail t…

Face GenerationTalking Face Generation

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

2026-04-26 · Zhen Ye, Xu Tan, Aoxiong Yin, Hongzhan Lin 외 arxiv

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attenti…

Video Generation

OmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style Replication

2025-04-03 · Zhongjian Wang, Peng Zhang, Jinwei Qi, Guangyuan Wang Sheng Xu 외

Recent years have witnessed remarkable advances in talking head generation, owing to its potential to revolutionize the human-AI interaction from text interfaces into realistic video chats. However, research on text-driv…

Talking Head GenerationVideo Synchronization

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

2025-06-01 · Zhengcong Fei, Hao Jiang, Di Qiu, Baoxuan Gu 외

The generation and editing of audio-conditioned talking portraits guided by multimodal inputs, including text, images, and videos, remains under explored. In this paper, we present SkyReels-Audio, a unified framework for…

Denoising

MakeItTalk: Speaker-Aware Talking-Head Animation

2020-04-27 · Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria 외

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or po…

Talking Face GenerationTalking Head Generation