paper-with-me

홈 › Papers

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

2025-11-05 · Guozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng, Youliang Zhang, Yi Chen, Yuan Zhou, Qinglin Lu, Limin Wang arxiv

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a unified framework for joint audio and video generation. UniAVGen is anchored in a dual-branch joint synthesis architecture, incorporating two parallel Diffusion Transformers (DiTs) to build a cohesive cross-modal latent space. At its heart lies an Asymmetric Cross-Modal Interaction mechanism, which enables bidirectional, temporally aligned cross-attention, thus ensuring precise spatiotemporal synchronization and semantic consistency. Furthermore, this cross-modal interaction is augmented by a Face-Aware Modulation module, which dynamically prioritizes salient regions in the interaction process. To enhance generative fidelity during inference, we additionally introduce Modality-Aware Classifier-Free Guidance, a novel strategy that explicitly amplifies cross-modal correlation signals. Notably, UniAVGen's robust joint synthesis design enables seamless unification of pivotal audio-video tasks within a single model, such as joint audio-video generation and continuation, video-to-audio dubbing, and audio-driven video synthesis. Comprehensive experiments validate that, with far fewer training samples (1.3M vs. 30.1M), UniAVGen delivers overall advantages in audio-video synchronization, timbre consistency, and emotion consistency.

📄 PDF Abstract BibTeX arXiv:2511.03334

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

LTX-2: Efficient Joint Audio-Visual Foundation Model

2026-01-06 · Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman 외 arxiv

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source found…

Audio GenerationVideo Generation

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

2026-06-29 · Kien T. Pham, I Chieh Chen, Qifeng Chen, Long Chen hf

Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained synchronization and semantic alignment between the auditory and visu…

Video ReconstructionVideo Generation

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

2025-02-06 · Lei Zhao, Linfeng Feng, Dongxu Ge, Rujin Chen 외

With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited exploration of unified generative architectures. …

Audio GenerationDiversityVideo Generation

Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation

2023-03-29 · Jiawei Liu, Weining Wang, Sihan Chen, Xinxin Zhu 외

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic vide…

Audio GenerationContrastive LearningDecoderVideo Generation

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation

2025-12-15 · Xiaohu Huang, Hao Zhou, Qiangpeng Yang, Shilei Wen 외 arxiv

In this paper, we present JoVA, a unified framework for joint video-audio generation. Despite recent encouraging advances, existing methods face two critical limitations. First, most existing approaches can only generate…

multimodal generationKeypoint DetectionAudio Generation