Perceptual Musical Features for Interpretable Audio Tagging
In the age of music streaming platforms, the task of automatically tagging music audio has garnered significant attention, driving researchers to devise methods aimed at enhancing performance metrics on standard datasets. Most recent approaches rely on deep neural networks, which, despite their impressive performance, possess opacity, making it challenging to elucidate their output for a given input. While the issue of interpretability has been emphasized in other fields like medicine, it has not received attention in music-related tasks. In this study, we explored the relevance of interpretability in the context of automatic music tagging. We constructed a workflow that incorporates three different information extraction techniques: a) leveraging symbolic knowledge, b) utilizing auxiliary deep neural networks, and c) employing signal processing to extract perceptual features from audio files. These features were subsequently used to train an interpretable machine-learning model for tag prediction. We conducted experiments on two datasets, namely the MTG-Jamendo dataset and the GTZAN dataset. Our method surpassed the performance of baseline models in both tasks and, in certain instances, demonstrated competitiveness with the current state-of-the-art. We conclude that there are use cases where the deterioration in performance is outweighed by the value of interpretability.
Code (1)
Tasks
Audio TaggingInterpretable Machine LearningMusic TaggingTAGSimilar Papers 제목 키워드 기반
A Modulation Front-End for Music Audio Tagging
Convolutional Neural Networks have been extensively explored in the task of automatic music tagging. The problem can be approached by using either engineered time-frequency features or raw audio as input. Modulation filt…
Audio TaggingMusic TaggingRepresentation LearningExpressivity-aware Music Performance Retrieval using Mid-level Perceptual Features and Emotion Word Embeddings
This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character,…
RetrievalWord EmbeddingsTwo-level Explanations in Music Emotion Recognition
Current ML models for music emotion recognition, while generally working quite well, do not give meaningful or intuitive explanations for their predictions. In this work, we propose a 2-step procedure to arrive at spectr…
Emotion RecognitionMusic Emotion RecognitionPredictionVocal Bursts Valence PredictionTracing Back Music Emotion Predictions to Sound Sources and Intuitive Perceptual Qualities
Music emotion recognition is an important task in MIR (Music Information Retrieval) research. Owing to factors like the subjective nature of the task and the variation of emotional cues between musical genres, there are …
Emotion RecognitionInformation RetrievalMusic Emotion RecognitionMusic Information Retrieval+2audioLIME: Listenable Explanations Using Source Separation
Deep neural networks (DNNs) are successfully applied in a wide variety of music information retrieval (MIR) tasks but their predictions are usually not interpretable. We propose audioLIME, a method based on Local Interpr…
Information RetrievalMusic Information RetrievalMusic TaggingRetrieval