paper-with-me

홈 › Papers

Inference-Time Scaling for Joint Audio-Video Generation

2026-06-02 · Jaemin Jung, Kyeongha Rho, Inkyu Shin, Joon Son Chung arxiv

Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint audio-video generation models often require substantial training resources to improve fidelity, Inference-Time Scaling (ITS) has recently emerged as a promising training-free alternative in single-modality domains. However, extending ITS from a single modality to multimodal domains is non-trivial, as it requires balancing multiple heterogeneous objectives. In this paper, we present the first comprehensive study of ITS for joint audio-video generation. We first demonstrate that a multi-verifier framework is essential to address the limitations of single-objective guidance, including asymmetric performance trade-offs and verifier hacking. Through systematic analysis, we then identify an optimal multi-verifier combination that yields balanced improvements across all quality dimensions. Finally, to effectively aggregate diverse reward signals, we propose Adaptive Reward Weighting (ARW), a novel test-time optimization algorithm. ARW treats reward aggregation as an online optimization problem, utilizing learnable parameters to calibrate reward variances without requiring prior knowledge of reward distributions, thereby ensuring robust multi-objective selection. Experimental results on VGGSound and JavisBench-mini benchmarks demonstrate that our framework significantly enhances semantic alignment, perceptual quality, and audio-visual synchronization of generated outputs. Synthesized samples and code are available on the project page: https://jung-jaemin.github.io/ITS-AVGen-Proj.

📄 PDF Abstract BibTeX arXiv:2606.03183

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

2026-04-21 · Liyang Li, Wen Wang, Canyu Zhao, Tianjian Feng 외 arxiv

Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within a single model. However, existing controllable generation framework…

Video Generation

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

2025-09-08 · Xiaoran Yang, Jianxuan Yang, Xinyue Guo, Haoyu Wang 외 arxiv

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instan…

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

2024-12-19 · CVPR 2025 1 · Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya 외

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (l…

Audio GenerationAudio SynthesisAudio-Visual SynchronizationVideo-to-Sound Generation

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

2025-11-18 · Keda Tao, Kele Shao, Bohan Yu, Weiqiang Wang 외 arxiv

Omnimodal large language models (OmniLLMs) have attracted increasing research attention of late towards unified audio-video understanding. However, the high computational cost of processing longer joint audio-video token…

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

2026-08-25 · Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu 외 arxiv

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. …

Audio Generation