paper-with-me

홈 › Papers

GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

2025-01-17 · Heda Zuo, Weitao You, Junxian Wu, Shihong Ren, Pei Chen, Mingxu Zhou, Yujia Lu, Lingyun Sun

Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative diversity, primarily due to inadequate feature alignment methods and insufficient datasets. In this study, we present General Video-to-Music Generation model (GVMGen), designed for generating high-related music to the video input. Our model employs hierarchical attentions to extract and align video features with music in both spatial and temporal dimensions, ensuring the preservation of pertinent features while minimizing redundancy. Remarkably, our method is versatile, capable of generating multi-style music from different video inputs, even in zero-shot scenarios. We also propose an evaluation model along with two novel objective metrics for assessing video-music alignment. Additionally, we have compiled a large-scale dataset comprising diverse types of video-music pairs. Experimental results demonstrate that GVMGen surpasses previous models in terms of music-video correspondence, generative diversity, and application universality.

📄 PDF Abstract BibTeX arXiv:2501.09972

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMusic Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation

2025-11-12 · Shulei Ji, Zihao Wang, Jiaxing Yu, Xiangyuan Yang 외 arxiv

Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignm…

Music Generation

Cross-modal Variational Auto-encoder for Content-based Micro-video Background Music Recommendation

2021-07-15 · Jing Yi, Yaochen Zhu, Jiayi Xie, Zhenzhong Chen

In this paper, we propose a cross-modal variational auto-encoder (CMVAE) for content-based micro-video background music recommendation. CMVAE is a hierarchical Bayesian generative model that matches relevant background m…

Music Recommendation

Vision-to-Music Generation: A Survey

2025-03-27 · Zhaokai Wang, Chenxi Bao, Le Zhuo, Jingrui Han 외

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short …

multimodal generationMusic GenerationSurvey

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

2023-05-11 · Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 외

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced au…

Music Generation

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

2026-07-10 · Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang 외 arxiv

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typi…