paper-with-me

Papers

MARBLE: Music Audio Representation Benchmark for Universal Evaluation

2023-06-18 · NeurIPS 2023 11 · Ruibin Yuan, Yinghao Ma, Yizhi Li, Ge Zhang, Xingran Chen, Hanzhi Yin, Le Zhuo, Yiqi Liu, Jiawen Huang, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger Dannenberg, Wenhu Chen, Gus Xia, Wei Xue, Si Liu, Shi Wang, Ruibo Liu, Yike Guo, Jie Fu

In the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 14 tasks on 8 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published at https://marble-bm.shef.ac.uk to promote future music AI research.

📄 PDF Abstract BibTeX arXiv:2306.10548

Code (1)

a43992899/marble-benchmark 공식 구현 pytorch

Tasks

Image GenerationInformation RetrievalMusic Information Retrieval

Similar Papers 제목 키워드 기반

USAD: Universal Speech and Audio Representation via Distillation

2025-06-23 · Heng-Jui Chang, Saurabhchand Bhati, James Glass, Alexander H. Liu

Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distill…

Audio TaggingRepresentation LearningSelf-Supervised LearningSound Classification

EnCodecMAE: Leveraging neural codecs for universal audio representation learning

2023-09-14 · Leonardo Pepino, Pablo Riera, Luciana Ferrer

The goal of universal audio representation learning is to obtain foundational models that can be used for a variety of downstream tasks involving speech, music and environmental sounds. To approach this problem, methods …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+2

UniWhisper: Efficient Continual Multi-task Training for Robust Universal Audio Representation

2026-02-25 · Yuxuan Chen, Peize He, Haoyuan Yu, Junzi Zhang arxiv

A universal audio representation should capture fine-grained speech cues and high-level semantics for environmental sounds and music in a single encoder. Existing encoders often excel in one domain but degrade in others.…

Universal Music Representations? Evaluating Foundation Models on World Music Corpora

2025-06-20 · Charilaos Papaioannou, Emmanouil Benetos, Alexandros Potamianos

Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of…

BenchmarkingFew-Shot LearningInformation RetrievalMusic Information Retrieval

High-Fidelity Music Vocoder using Neural Audio Codecs

2025-02-18 · Luca A. Lanzendörfer, Florian Grötschla, Michael Ungersböck, Roger Wattenhofer

While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages…

DecoderSpeech Synthesis