Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio
We present PECMAE, an interpretable model for music audio classification based on prototype learning. Our model is based on a previous method, APNet, which jointly learns an autoencoder and a prototypical network. Instead, we propose to decouple both training processes. This enables us to leverage existing self-supervised autoencoders pre-trained on much larger data (EnCodecMAE), providing representations with better generalization. APNet allows prototypes' reconstruction to waveforms for interpretability relying on the nearest training data samples. In contrast, we explore using a diffusion decoder that allows reconstruction without such dependency. We evaluate our method on datasets for music instrument classification (Medley-Solos-DB) and genre recognition (GTZAN and a larger in-house dataset), the latter being a more challenging task not addressed with prototypical networks before. We find that the prototype-based models preserve most of the performance achieved with the autoencoder embeddings, while the sonification of prototypes benefits understanding the behavior of the classifier.
Code (1)
Tasks
Audio ClassificationDecoderMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Style-Aware Symbolic Music Representations by Adversarial Autoencoders
We address the challenging open problem of learning an effective latent space for symbolic music data in generative music modeling. We focus on leveraging adversarial regularization as a flexible and natural mean to imbu…
Music ModelingLearning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature requires compression that obscures semant…
Audio GenerationMusic GenerationAdvancing Cultural Inclusivity: Optimizing Embedding Spaces for Balanced Music Recommendations
Popularity bias in music recommendation systems -- where artists and tracks with the highest listen counts are recommended more often -- can also propagate biases along demographic and cultural axes. In this work, we ide…
FairnessMusic RecommendationRecommendation SystemsMusic2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
Efficiently compressing high-dimensional audio signals into a compact and informative latent space is crucial for various tasks, including generative modeling and music information retrieval (MIR). Existing audio autoenc…
Audio CompressionDenoisingInformation RetrievalMusic Information Retrieval+1Music2Latent: Consistency Autoencoders for Latent Audio Compression
Efficient audio representations in a compressed continuous latent space are critical for generative audio modeling and Music Information Retrieval (MIR) tasks. However, some existing audio autoencoders have limitations, …
Audio CompressionInformation RetrievalMusic Information Retrieval