paper-with-me

홈 › Papers

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis

2026-03-28 · Axel Marmoret arxiv

Music Structure Analysis (MSA) aims to uncover the high-level organization of musical pieces. State-of-the-art methods are often based on supervised deep learning, but these methods are bottlenecked by the need for heavily annotated data and inherent structural ambiguities. In this paper, we propose an unsupervised evaluation of nine open-source, generic pre-trained deep audio models, on MSA. For each model, we extract barwise embeddings and segment them using three unsupervised segmentation algorithms (Foote's checkerboard kernels, spectral clustering, and Correlation Block-Matching (CBM)), focusing exclusively on boundary retrieval. Our results demonstrate that modern, generic deep embeddings generally outperform traditional spectrogram-based baselines, but not systematically. Furthermore, our unsupervised boundary estimation methodology generally yields stronger performance than recent linear probing baselines. Among the evaluated techniques, the CBM algorithm consistently emerges as the most effective downstream segmentation method. Finally, we highlight the artificial inflation of standard evaluation metrics and advocate for the systematic adoption of `trimming'', or even `double trimming'' annotations to establish more rigorous MSA evaluation standards.

📄 PDF Abstract BibTeX arXiv:2603.27218

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unsupervised Learning of Deep Features for Music Segmentation

2021-08-30 · Matthew C. McCallum

Music segmentation refers to the dual problem of identifying boundaries between, and labeling, distinct music segments, e.g., the chorus, verse, bridge etc. in popular music. The performance of a range of music segmentat…

SegmentationSound Classification

Learning Normal Patterns in Musical Loops

2025-05-22 · Shayan Dadman, Bernt Arild Bremdal, Børre Bang, Rune Dalmo

This paper introduces an unsupervised framework for detecting audio patterns in musical samples (loops) through anomaly detection techniques, addressing challenges in music information retrieval (MIR). Existing methods a…

Anomaly DetectionInformation RetrievalMusic Information RetrievalUnsupervised Anomaly Detection

Adapting Frechet Audio Distance for Generative Music Evaluation

2023-11-02 · Azalea Gui, Hannes Gamper, Sebastian Braun, Dimitra Emmanouilidou

The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correla…

FAD

Supervised and Unsupervised Learning of Audio Representations for Music Understanding

2022-10-07 · Matthew C. McCallum, Filip Korzeniowski, Sergio Oramas, Fabien Gouyon 외

In this work, we provide a broad comparative analysis of strategies for pre-training audio understanding models for several tasks in the music domain, including labelling of genre, era, origin, mood, instrumentation, key…

Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation

2024-12-10 · Jan Retkowski, Jakub Stępniak, Mateusz Modrzejewski

In this paper we introduce the Frechet Music Distance (FMD), a novel evaluation metric for generative symbolic music models, inspired by the Frechet Inception Distance (FID) in computer vision and Frechet Audio Distance …

FADMusic GenerationMusic Modeling