Inspecting and Interacting with Meaningful Music Representations using VAE
Variational Autoencoders(VAEs) have already achieved great results on image generation and recently made promising progress on music generation. However, the generation process is still quite difficult to control in the sense that the learned latent representations lack meaningful music semantics. It would be much more useful if people can modify certain music features, such as rhythm and pitch contour, via latent representations to test different composition ideas. In this paper, we propose a new method to inspect the pitch and rhythm interpretations of the latent representations and we name it disentanglement by augmentation. Based on the interpretable representations, an intuitive graphical user interface is designed for users to better direct the music creation process by manipulating the pitch contours and rhythmic complexity.
Code (0)
등록된 구현이 없습니다.
Tasks
DisentanglementImage GenerationMusic GenerationRhythmSimilar Papers 제목 키워드 기반
Melodic Contour and Mid-Level Global Features Applied to the Analysis of Flamenco Cantes
This work focuses on the topic of melodic characterization and similarity in a specific musical repertoire: a cappella flamenco singing, more specifically in debla and martinete styles. We propose the combination of manu…
S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification
In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S…
ClassificationData AugmentationGenre classificationMusic Classification+3Are Nearby Neighbors Relatives?: Testing Deep Music Embeddings
Deep neural networks have frequently been used to directly learn representations useful for a given task from raw input data. In terms of overall performance metrics, machine learning solutions employing deep representat…
GraphMuse: A Library for Symbolic Music Graph Processing
Graph Neural Networks (GNNs) have recently gained traction in symbolic music tasks, yet a lack of a unified framework impedes progress. Addressing this gap, we present GraphMuse, a graph processing framework and library …
Cross-Modal Learning for Music-to-Music-Video Description Generation
Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models has opened a promising pathway for musi…
Video DescriptionVideo Generation