paper-with-me

Papers

Do Foundational Audio Encoders Understand Music Structure?

2025-12-19 · Keisuke Toyama, Zhi Zhong, Akira Takahashi, Shusuke Takahashi, Yuki Mitsufuji arxiv

In music information retrieval (MIR) research, the use of pretrained foundational audio encoders (FAEs) has recently become a trend. FAEs pretrained on large amounts of music and audio data have been shown to improve performance on MIR tasks such as music tagging and automatic music transcription. However, their use for music structure analysis (MSA) remains underexplored: only a small subset of FAEs has been examined for MSA, and the impact of factors such as learning methods, training data, and model context length on MSA performance remains unclear. In this study, we conduct comprehensive experiments on 11 types of FAEs to investigate how these factors affect MSA performance. Our results demonstrate that FAEs using self-supervised learning with masked language modeling on music data are particularly effective for MSA. These findings pave the way for future research in FAE and MSA.

📄 PDF Abstract BibTeX arXiv:2512.17209

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningInformation RetrievalMusic Transcription

Similar Papers 제목 키워드 기반

Music Flamingo: Scaling Music Understanding in Audio Language Models

2025-11-13 · Sreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee 외 arxiv

We introduce Music Flamingo, a novel large audio-language model designed to advance music (including song) understanding in foundational audio models. While audio-language research has progressed rapidly, music remains c…

Reinforcement Learning

From Audio Encoders to Piano Judges: Benchmarking Performance Understanding for Solo Piano

2024-07-05 · huan zhang, Jinhua Liang, Simon Dixon

Our study investigates an approach for understanding musical performances through the lens of audio encoding models, focusing on the domain of solo Western classical piano music. Compared to composition-level attribute u…

AttributeBenchmarking

EnCodecMAE: Leveraging neural codecs for universal audio representation learning

2023-09-14 · Leonardo Pepino, Pablo Riera, Luciana Ferrer

The goal of universal audio representation learning is to obtain foundational models that can be used for a variety of downstream tasks involving speech, music and environmental sounds. To approach this problem, methods …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+2

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

2026-06-04 · Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati, Mrudula Athi 외 arxiv

Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. While self-supervised learning (SSL) has yielded strong domain-specific e…

Self-Supervised Learning

Exploring Musical Roots: Applying Audio Embeddings to Empower Influence Attribution for a Generative Music Model

2024-01-25 · Julia Barnett, Hugo Flores Garcia, Bryan Pardo

Every artist has a creative process that draws inspiration from previous artists and their works. Today, "inspiration" has been automated by generative music models. The black box nature of these models obscures the iden…