paper-with-me

Papers

Learning music audio representations via weak language supervision

2021-12-08 · Ilaria Manco, Emmanouil Benetos, Elio Quinton, Gyorgy Fazekas

Audio representations for music information retrieval are typically learned via supervised learning in a task-specific fashion. Although effective at producing state-of-the-art results, this scheme lacks flexibility with respect to the range of applications a model can have and requires extensively annotated datasets. In this work, we pose the question of whether it may be possible to exploit weakly aligned text as the only supervisory signal to learn general-purpose music audio representations. To address this question, we design a multimodal architecture for music and language pre-training (MuLaP) optimised via a set of proxy tasks. Weak supervision is provided in the form of noisy natural language descriptions conveying the overall musical content of the track. After pre-training, we transfer the audio backbone of the model to a set of music audio classification and regression tasks. We demonstrate the usefulness of our approach by comparing the performance of audio representations produced by the same audio backbone with different training strategies and show that our pre-training method consistently achieves comparable or higher scores on all tasks and datasets considered. Our experiments also confirm that MuLaP effectively leverages audio-caption pairs to learn representations that are competitive with audio-only and cross-modal self-supervised methods in the literature.

📄 PDF Abstract BibTeX arXiv:2112.04214

Code (1)

ilaria-manco/mulap 공식 구현 pytorch

Tasks

Audio ClassificationInformation RetrievalMusic Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

2025-05-20 · Ziqian Wang, Xianjun Xia, Xinfa Zhu, Lei Xie

The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse au…

cross-modal alignmentLanguage ModelingLanguage ModellingLarge Language Model+2

Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

2026-03-26 · Xuanru Zhou, Yiwen Shao, Wei-Cheng Tseng, Dong Yu arxiv

Current audio pre-training seeks to learn unified representations for broad audio understanding tasks, but it remains fragmented and is fundamentally bottlenecked by its reliance on weak, noisy, and scale-limited labels.…

Count The Notes: Histogram-Based Supervision for Automatic Music Transcription

2025-11-18 · Jonathan Yaffe, Ben Maman, Meinard Müller, Amit H. Bermano arxiv

Automatic Music Transcription (AMT) converts audio recordings into symbolic musical representations. Training deep neural networks (DNNs) for AMT typically requires strongly aligned training pairs with precise frame-leve…

Music Transcription

The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS

2025-10-21 · Brandon James Carone, Iran R. Roman, Pablo Ripollés arxiv

Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and…

Relational Reasoning

Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision

2018-11-09 · Sanjeel Parekh, Alexey Ozerov, Slim Essid, Ngoc Duong 외

We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object classification in noisy acoustic enviro…

General ClassificationMultiple Instance LearningObjectObject Localization+1