paper-with-me

홈 › Papers

MAVIN: Multi-Shot Audio-Visual Generation with Narrative Control

2026-06-28 · Kaiqi Liu, Yunyao Mao, Ziqi Cai, Zheng Geng, Jing Wang, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, Shuchen Weng, Boxin Shi arxiv

While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-visual generation. Existing methods suffer from temporal misalignment, limited controllability, and incomplete scripting. In this paper, we propose MAVIN, the first framework for multi-shot audio-visual generation with customized narrative control. To resolve temporal misalignment, we propose boundary-aware attention, which leverages hierarchical captions and boundary-aware token routing to render audio-visual elements within their respective temporal boundaries. To improve the controllability for multi-subject scenarios, we propose ID-aware propagation, utilizing identity embeddings and an identity-aware mask to bind specific identities to consistent visual appearances and vocal timbres. To provide comprehensive audio-visual narratives, we present a multi-agent scripting pipeline to transform free-form user inputs into hierarchical captions. Furthermore, we construct MAVINSet, a multi-shot audio-visual dataset for robust training and evaluation. Extensive experiments demonstrate that MAVIN achieves state-of-the-art performance, opening up a new avenue for integrating generative models into professional filmmaking workflows.

📄 PDF Abstract BibTeX arXiv:2606.29473

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAVIN: Multi-Action Video Generation with Diffusion Models via Transition Video Infilling

2024-05-28 · BoWen Zhang, Xiaofei Xie, Haotian Lu, Na Ma 외

Diffusion-based video generation has achieved significant progress, yet generating multiple actions that occur sequentially remains a formidable task. Directly generating a video with sequential actions can be extremely …

Video Generation

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

2024-12-12 · Zihao Chen, Haomin Zhang, Xinhan Di, Haoyu Wang 외

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge o…

Audio Generation

SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound

2024-06-06 · Rishit Dagli, Shivesh Prakash, Robert Wu, Houman Khosravani

Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across mu…

Audio Generation

MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation

2025-03-14 · Sungwoo Cho, Jeongsoo Choi, Sungnyun Kim, Se-Young Yun

Despite recent advances in text-to-speech (TTS) models, audio-visual to audio-visual (AV2AV) translation still faces a critical challenge: maintaining speaker consistency between the original and translated vocal and fac…

text-to-speechText to SpeechTranslation

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

2026-06-19 · Jiehui Huang, Yuechen Zhang, Bin Xia, Jiahao Wang 외 arxiv

Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches either train end-to-end over fixed-lengt…

Video Generation