paper-with-me

Papers

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

2026-04-26 · Zhen Ye, Xu Tan, Aoxiong Yin, Hongzhan Lin, Guangyan Zhang, Peiwen Sun, Yiming Li, Chi-Min Chan, Wei Ye, Shikun Zhang, Wei Xue arxiv

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating high-level semantics and low-level details in a fully entangled manner. This is suboptimal for talking head synthesis: while audio and facial motion are semantically correlated, their low-level realizations (acoustic signals and visual textures) follow distinct rendering processes. Enforcing joint modeling across all levels causes unnecessary entanglement and reduces efficiency. We propose Talker-T2AV, an autoregressive diffusion framework where high-level cross-modal modeling occurs in a shared backbone, while low-level refinement uses modality-specific decoders. A shared autoregressive language model jointly reasons over audio and video in a unified patch-level token space. Two lightweight diffusion transformer heads decode the hidden states into frame-level audio and video latents. Experiments on talking portrait benchmarks show Talker-T2AV outperforms dual-branch baselines in lip-sync accuracy, video quality, and audio quality, achieving stronger cross-modal consistency than cascaded pipelines.

📄 PDF Abstract BibTeX arXiv:2604.23586

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

2022-08-23 · Dongchan Min, Minyoung Song, Eunji Ko, Sung Ju Hwang

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and …

Talking Head GenerationVideo Generation

OmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style Replication

2025-04-03 · Zhongjian Wang, Peng Zhang, Jinwei Qi, Guangyuan Wang Sheng Xu 외

Recent years have witnessed remarkable advances in talking head generation, owing to its potential to revolutionize the human-AI interaction from text interfaces into realistic video chats. However, research on text-driv…

Talking Head GenerationVideo Synchronization

GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting

2024-04-29 · Bo Chen, Shoukang Hu, Qi Chen, Chenpeng Du 외

We present GStalker, a 3D audio-driven talking face generation model with Gaussian Splatting for both fast training (40 minutes) and real-time rendering (125 FPS) with a 3$\sim$5 minute video for training material, in co…

Face GenerationNeRFTalking Face Generation

EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation

2025-10-03 · Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of …

Talking Head Generation

NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis

2025-02-20 · Xiaoxing Liu, Zhilei Liu, Chongke Bi

Talking head synthesis is to synthesize a lip-synchronized talking head video using audio. Recently, the capability of NeRF to enhance the realism and texture details of synthesized talking heads has attracted the attent…

DisentanglementNeRF