paper-with-me

홈 › Papers

From Sound to Sight: Towards AI-authored Music Videos

2025-08-20 · Leo Vitasovic, Stella Graßhof, Agnes Mercedes Kloft, Ville V. Lehtola, Martin Cunneen, Justyna Starostka, Glenn McGarry, Kun Li, Sami S. Brandt arxiv

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any user-specified, vocal or instrumental song using off-the-shelf deep learning models. Inspired by the manual workflows of music video producers, we experiment on how well latent feature-based techniques can analyse audio to detect musical qualities, such as emotional cues and instrumental patterns, and distil them into textual scene descriptions using a language model. Next, we employ a generative model to produce the corresponding video clips. To assess the generated videos, we identify several critical aspects and design and conduct a preliminary user evaluation that demonstrates storytelling potential, visual coherency and emotional alignment with the music. Our findings underscore the potential of latent feature techniques and deep generative models to expand music visualisation beyond traditional approaches.

📄 PDF Abstract BibTeX arXiv:2509.00029

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InverseMV: Composing Piano Scores with a Convolutional Video-Music Transformer

2021-12-31 · Chin-Tung Lin, Mu Yang

Many social media users prefer consuming content in the form of videos rather than text. However, in order for content creators to produce videos with a high click-through rate, much editing is needed to match the footag…

Music Generation

How Does it Sound?

2021-12-01 · NeurIPS 2021 12 · Kun Su, Xiulong Liu, Eli Shlizerman

One of the primary purposes of video is to capture people and their unique activities. It is often the case that the experience of watching the video can be enhanced by adding a musical soundtrack that is in-sync with th…

Rhythm

Musical Audio Similarity with Self-supervised Convolutional Neural Networks

2022-02-04 · Carl Thomé, Sebastian Piwell, Oscar Utterbäck

We have built a music similarity search engine that lets video producers search by listenable music excerpts, as a complement to traditional full-text search. Our system suggests similar sounding track segments in a larg…

Triplet

Automated Composition of Picture-Synched Music Soundtracks for Movies

2019-10-19 · Vansh Dassani, Jon Bird, Dave Cliff

We describe the implementation of and early results from a system that automatically composes picture-synched musical soundtracks for videos and movies. We use the phrase "picture-synched" to mean that the structure of t…

Music Generation

The NES Video-Music Database: A Dataset of Symbolic Video Game Music Paired with Gameplay Videos

2024-04-05 · Igor Cardoso, Rubens O. Moraes, Lucas N. Ferreira

Neural models are one of the most popular approaches for music generation, yet there aren't standard large datasets tailored for learning music directly from game data. To address this research gap, we introduce a novel …

Music Generation