paper-with-me

홈 › Papers

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

2026-02-22 · Kai Liu, Yanhao Zheng, Kai Wang, Shengqiong Wu, Rongjunchen Zhang, Jiebo Luo, Dimitrios Hatzinakos, Ziwei Liu, Hao Fei, Tat-Seng Chua arxiv

AIGC has rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision from textual descriptions. However, compared with advanced commercial models such as Veo3, existing open-source methods still suffer from limitations in generation quality, temporal synchrony, and alignment with human preferences. To bridge the gap, this paper presents JavisDiT++, a concise yet powerful framework for unified modeling and optimization of JAVG. First, we introduce a modality-specific mixture-of-experts (MS-MoE) design that enables cross-modal interaction efficacy while enhancing single-modal generation quality. Then, we propose a temporal-aligned RoPE (TA-RoPE) strategy to achieve explicit, frame-level synchronization between audio and video tokens. Besides, we develop an audio-video direct preference optimization (AV-DPO) method to align model outputs with human preference across quality, consistency, and synchrony dimensions. Built upon Wan2.1-1.3B-T2V, our model achieves state-of-the-art performance merely with around 1M public training entries, significantly outperforming prior approaches in both qualitative and quantitative evaluations. Comprehensive ablation studies have been conducted to validate the effectiveness of our proposed modules. All the code, model, and dataset are released at https://JavisVerse.github.io/JavisDiT2-page.

📄 PDF Abstract BibTeX arXiv:2602.19163

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image GenerationVideo Generation

Similar Papers 제목 키워드 기반

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

2025-03-30 · Kai Liu, Wei Li, Lai Chen, Shengqiong Wu 외

This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Built upon the powerful Diffusion Transformer (DiT) architecture, JavisDiT is able …

Video Generation

Apollo: Unified Multi-Task Audio-Video Joint Generation

2026-01-07 · Jun Wang, Chunyu Qiang, Yuxin Guo, Yiran Wang 외 arxiv

Audio-video joint generation has progressed rapidly, yet substantial challenges still remain. Non-commercial approaches still suffer audio-visual asynchrony, poor lip-speech alignment, and unimodal degradation, which can…

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

2026-04-26 · Zhen Ye, Xu Tan, Aoxiong Yin, Hongzhan Lin 외 arxiv

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attenti…

Video Generation

InstructAudio: Unified speech and music generation with natural language instruction

2025-11-23 · Chunyu Qiang, Kang Yin, Xiaopeng Wang, Yuzhe Liang 외 arxiv

Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, an…

Multi-Task LearningDialogue GenerationMusic GenerationSpeech Synthesis

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

2026-08-31 · Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang 외 hf

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video genera…

Reinforcement LearningVideo Generation