paper-with-me

Papers

Video-based Music Generation

2026-02-05 · Serkan Sulun arxiv

As the volume of video content on the internet grows rapidly, finding a suitable soundtrack remains a significant challenge. This thesis presents EMSYNC (EMotion and SYNChronization), a fast, free, and automatic solution that generates music tailored to the input video, enabling content creators to enhance their productions without composing or licensing music. Our model creates music that is emotionally and rhythmically synchronized with the video. A core component of EMSYNC is a novel video emotion classifier. By leveraging pretrained deep neural networks for feature extraction and keeping them frozen while training only fusion layers, we reduce computational complexity while improving accuracy. We show the generalization abilities of our method by obtaining state-of-the-art results on Ekman-6 and MovieNet. Another key contribution is a large-scale, emotion-labeled MIDI dataset for affective music generation. We then present an emotion-based MIDI generator, the first to condition on continuous emotional values rather than discrete categories, enabling nuanced music generation aligned with complex emotional content. To enhance temporal synchronization, we introduce a novel temporal boundary conditioning method, called "boundary offset encodings," aligning musical chords with scene changes. Combining video emotion classification, emotion-based music generation, and temporal boundary conditioning, EMSYNC emerges as a fully automatic video-based music generator. User studies show that it consistently outperforms existing methods in terms of music richness, emotional alignment, temporal synchronization, and overall preference, setting a new state-of-the-art in video-based music generation.

📄 PDF Abstract BibTeX arXiv:2602.07063

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion ClassificationMusic Generation

Similar Papers 제목 키워드 기반

Video Background Music Generation: Dataset, Method and Evaluation

2022-11-21 · ICCV 2023 1 · Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao 외

Music is essential when editing videos, but selecting music manually is difficult and time-consuming. Thus, we seek to automatically generate background music tracks given video input. This is a challenging task since it…

Music GenerationRepresentation LearningRetrieval

VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

2024-09-11 · Yan-Bo Lin, Yu Tian, Linjie Yang, Gedas Bertasius 외

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large…

Contrastive LearningMusic Generation

Cross-Modal Learning for Music-to-Music-Video Description Generation

2025-03-14 · Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Zhi Zhong 외

Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models has opened a promising pathway for musi…

Video DescriptionVideo Generation

Diff-BGM: A Diffusion Model for Video Background Music Generation

2024-05-20 · CVPR 2024 1 · Sizhe Li, Yiming Qin, Minghang Zheng, Xin Jin 외

When editing a video, a piece of attractive background music is indispensable. However, video background music generation tasks face several challenges, for example, the lack of suitable training datasets, and the diffic…

DiversityMusic GenerationRhythm

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

2023-05-11 · Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 외

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced au…

Music Generation