paper-with-me

Papers

V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

2026-03-11 · Yan-Bo Lin, Jonah Casebeer, Long Mai, Aniruddha Mahapatra, Gedas Bertasius, Nicholas J. Bryan arxiv

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned music with disentangled time synchronization and semantic control (e.g., genre, mood) from video while requiring zero video-music pairs at training time. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra-modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross-modal training or paired data. Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-ZERO achieves state-of-the-art performance without any paired music-video data, surpassing the strongest prior baselines per metric with 5-9% higher audio quality, 13-15% better semantic alignment, 21-52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Our results validate that temporal alignment through within-modality features is not only effective for video-to-music generation but also leads to better performance than paired cross-modal supervision. Furthermore, our approach enables independent controls for timing and music style (e.g., genre, mood) for more controllable generation.

📄 PDF Abstract BibTeX arXiv:2603.11042

Code (0)

등록된 구현이 없습니다.

Tasks

Music Generation

Similar Papers 제목 키워드 기반

Video-Zero: Self-Evolution Video Understanding

2026-05-14 · Ruixu Zhang, Deyi Ji, Lanyun Zhu, Xuanyi Liu 외 arxiv

Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to video understanding remains underexplored and challenging: videos a…

Analyzing Zero-Shot Abilities of Vision-Language Models on Video Understanding Tasks

2023-10-07 · Avinash Madasu, Anahita Bhiwandiwalla, Vasudev Lal

Foundational multimodal models pre-trained on large scale image-text pairs or video-text pairs or both have shown strong generalization abilities on downstream tasks. However unlike image-text models, pretraining video-t…

Action RecognitionMultiple-choiceQuestion AnsweringTemporal Action Localization+4

Semantically Tied Paired Cycle Consistency for Zero-Shot Sketch-based Image Retrieval

2019-03-08 · CVPR 2019 6 · Anjan Dutta, Zeynep Akata

Zero-shot sketch-based image retrieval (SBIR) is an emerging task in computer vision, allowing to retrieve natural images relevant to sketch queries that might not been seen in the training phase. Existing works either r…

feature selectionImage RetrievalRetrievalSketch-Based Image Retrieval

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner

2024-12-13 · Yufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Nanxuan Zhao 외

We present SUGAR, a zero-shot method for subject-driven video customization. Given an input image, SUGAR is capable of generating videos for the subject contained in the image and aligning the generation with arbitrary v…

Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors

2024-11-26 · CVPR 2025 1 · Zhengfei Kuang, Tianyuan Zhang, Kai Zhang, Hao Tan 외

We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and video--normal training data. Instead of rel…

Optical Flow Estimation