Vis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation
In this study, we explore the representation mapping from the domain of visual arts to the domain of music, with which we can use visual arts as an effective handle to control music generation. Unlike most studies in multimodal representation learning that are purely data-driven, we adopt an analysis-by-synthesis approach that combines deep music representation learning with user studies. Such an approach enables us to discover \textit{interpretable} representation mapping without a huge amount of paired data. In particular, we discover that visual-to-music mapping has a nice property similar to equivariant. In other words, we can use various image transformations, say, changing brightness, changing contrast, style transfer, to control the corresponding transformations in the music domain. In addition, we released the Vis2Mus system as a controllable interface for symbolic music generation.
Code (1)
Tasks
Music GenerationRepresentation LearningStyle TransferSimilar Papers 제목 키워드 기반
Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation
Multimodal music generation aims to produce music from diverse input modalities, including text, videos, and images. Existing methods use a common embedding space for multimodal fusion. Despite their effectiveness in oth…
cross-modal alignmentMultimodal Music GenerationMusic GenerationRetrievalLanguage Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal
We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappi…
cross-modal alignmentLanguage ModelingLanguage ModellingRepresentation LearningExploring Softly Masked Language Modelling for Controllable Symbolic Music Generation
This document presents some early explorations of applying Softly Masked Language Modelling (SMLM) to symbolic music generation. SMLM can be seen as a generalisation of masked language modelling (MLM), where instead of e…
Language ModellingMusic GenerationOpenDance: Multimodal Controllable 3D Dance Generation Using Large-scale Internet Data
Music-driven dance generation offers significant creative potential yet faces considerable challenges. The absence of fine-grained multimodal data and the difficulty of flexible multi-conditional generation limit previou…
DiversityEvaluating Disentangled Representations for Controllable Music Generation
Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings r…
Music Generation