paper-with-me

Papers

MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

2024-10-16 · RuiQi Li, Siqi Zheng, Xize Cheng, Ziang Zhang, Shengpeng Ji, Zhou Zhao

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual narratives. This paper presents MuVi, a novel framework that effectively addresses these challenges to enhance the cohesion and immersive experience of audio-visual content. MuVi analyzes video content through a specially designed visual adaptor to extract contextually and temporally relevant features. These features are used to generate music that not only matches the video's mood and theme but also its rhythm and pacing. We also introduce a contrastive music-visual pre-training scheme to ensure synchronization, based on the periodicity nature of music phrases. In addition, we demonstrate that our flow-matching-based music generator has in-context learning ability, allowing us to control the style and genre of the generated music. Experimental results show that MuVi demonstrates superior performance in both audio quality and temporal synchronization. The generated music video samples are available at https://muvi-v2m.github.io.

📄 PDF Abstract BibTeX arXiv:2410.12957

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningMusic GenerationRhythm

Similar Papers 제목 키워드 기반

Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model

2023-11-02 · Jaeyong Kang, Soujanya Poria, Dorien Herremans

Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos. In this work, we develop a generative …

Music GenerationRhythm

VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

2024-09-11 · Yan-Bo Lin, Yu Tian, Linjie Yang, Gedas Bertasius 외

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large…

Contrastive LearningMusic Generation

V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

2026-03-11 · Yan-Bo Lin, Jonah Casebeer, Long Mai, Aniruddha Mahapatra 외 arxiv

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that gener…

Music Generation

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

2026-04-19 · Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj, Gouthaman KV 외 arxiv

Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditioning alone and provide limited semantic…

Music Generation

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks

2025-05-26 · Chang Liu, Haomin Zhang, Shiyu Xia, Zihao Chen 외

Generating high-quality piano audio from video requires precise synchronization between visual cues and musical output, ensuring accurate semantic and temporal alignment.However, existing evaluation datasets do not fully…

Music Generation