Zero-shot Learning for Audio-based Music Classification and Tagging
Audio-based music classification and tagging is typically based on categorical supervised learning with a fixed set of labels. This intrinsically cannot handle unseen labels such as newly added music genres or semantic words that users arbitrarily choose for music retrieval. Zero-shot learning can address this problem by leveraging an additional semantic space of labels where side information about the labels is used to unveil the relationship between each other. In this work, we investigate the zero-shot learning in the music domain and organize two different setups of side information. One is using human-labeled attribute information based on Free Music Archive and OpenMIC-2018 datasets. The other is using general word semantic information based on Million Song Dataset and Last.fm tag annotations. Considering a music track is usually multi-labeled in music classification and tagging datasets, we also propose a data split scheme and associated evaluation settings for the multi-label zero-shot learning. Finally, we report experimental results and discuss the effectiveness and new possibilities of zero-shot learning in the music domain.
Code (1)
Tasks
AttributeClassificationGeneral ClassificationMulti-label zero-shot learningMusic ClassificationRetrievalTAGZero-Shot LearningSimilar Papers 제목 키워드 기반
Zero-shot Learning and Knowledge Transfer in Music Classification and Tagging
Music classification and tagging is conducted through categorical supervised learning with a fixed set of labels. In principle, this cannot make predictions on unseen labels. Zero-shot learning is an approach to solve th…
ClassificationGeneral ClassificationMusic ClassificationTransfer Learning+1Joint Music and Language Attention Models for Zero-shot Music Tagging
Music tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be generalized to new tags. In this work, we prop…
Audio TaggingDecoderMusic TaggingMuLan: A Joint Embedding of Music Audio and Natural Language
Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This paper presents MuLan: a first attempt at a…
Cross-Modal RetrievalMusic TaggingRetrievalTransfer LearningContrastive Audio-Language Learning for Music
As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information R…
Audio to Text RetrievalDescriptiveGenre classificationInformation Retrieval+3LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
We introduce Label-Combination Prototypical Networks (LC-Protonets) to address the problem of multi-label few-shot classification, where a model must generalize to new classes based on only a few available examples. Exte…
Audio TaggingFew-Shot LearningMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION