Metric Learning vs Classification for Disentangled Music Representation Learning
Deep representation learning offers a powerful paradigm for mapping input data onto an organized embedding space and is useful for many music information retrieval tasks. Two central methods for representation learning include deep metric learning and classification, both having the same goal of learning a representation that can generalize well across tasks. Along with generalization, the emerging concept of disentangled representations is also of great interest, where multiple semantic concepts (e.g., genre, mood, instrumentation) are learned jointly but remain separable in the learned representation space. In this paper we present a single representation learning framework that elucidates the relationship between metric learning, classification, and disentanglement in a holistic manner. For this, we (1) outline past work on the relationship between metric learning and classification, (2) extend this relationship to multi-label data by exploring three different learning approaches and their disentangled versions, and (3) evaluate all models on four tasks (training time, similarity retrieval, auto-tagging, and triplet prediction). We find that classification-based models are generally advantageous for training time, similarity retrieval, and auto-tagging, while deep metric learning exhibits better performance for triplet-prediction. Finally, we show that our proposed approach yields state-of-the-art results for music auto-tagging.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationDisentanglementGeneral ClassificationInformation RetrievalMetric LearningMusic Auto-TaggingMusic Information RetrievalRepresentation LearningRetrievalTripletSimilar Papers 제목 키워드 기반
Evaluating Disentangled Representations for Controllable Music Generation
Recent approaches in music generation rely on disentangled representations, often labeled as structure and timbre or local and global, to enable controllable synthesis. Yet the underlying properties of these embeddings r…
Music GenerationMulti-input Architecture and Disentangled Representation Learning for Multi-dimensional Modeling of Music Similarity
In the context of music information retrieval, similarity-based approaches are useful for a variety of tasks that benefit from a query-by-example scenario. Music however, naturally decomposes into a set of semantically m…
DisentanglementInformation RetrievalMusic Information RetrievalRepresentation Learning+2Signal-domain representation of symbolic music for learning embedding spaces
A key aspect of machine learning models lies in their ability to learn efficient intermediate features. However, the input representation plays a crucial role in this process, and polyphonic musical scores remain a parti…
Disentangled Multidimensional Metric Learning for Music Similarity
Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessa…
Metric LearningSpecificityVideo EditingSelf-Supervised VQ-VAE for One-Shot Music Style Transfer
Neural style transfer, allowing to apply the artistic style of one image to another, has become one of the most widely showcased computer vision applications shortly after its introduction. In contrast, related tasks in …
Music Style TransferSelf-Supervised LearningStyle Transfer