paper-with-me

Papers

Multi-Task Self-Supervised Pre-Training for Music Classification

2021-02-05 · Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang, Ming Sun, Brian McFee, Juan Pablo Bello, Chao Wang

Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. Besides, models learned from labeled dataset often embed biases specific to that particular dataset. Therefore, unsupervised learning techniques become popular approaches in solving machine listening problems. Particularly, a self-supervised learning technique utilizing reconstructions of multiple hand-crafted audio features has shown promising results when it is applied to speech domain such as emotion recognition and automatic speech recognition (ASR). In this paper, we apply self-supervised and multi-task learning methods for pre-training music encoders, and explore various design choices including encoder architectures, weighting mechanisms to combine losses from multiple tasks, and worker selections of pretext tasks. We investigate how these design choices interact with various downstream music classification tasks. We find that using various music specific workers altogether with weighting mechanisms to balance the losses during pre-training helps improve and generalize to the downstream tasks.

📄 PDF Abstract BibTeX arXiv:2102.03229

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationEmotion RecognitionGeneral ClassificationMulti-Task LearningMusic ClassificationSelf-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification

2022-02-21 · Hang Zhao, Chen Zhang, Belei Zhu, Zejun Ma 외

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S…

ClassificationData AugmentationGenre classificationMusic Classification+3

MusiCoder: A Universal Music-Acoustic Encoder Based on Transformers

2020-08-03 · Yilun Zhao, Jia Guo

Music annotation has always been one of the critical topics in the field of Music Information Retrieval (MIR). Traditional models use supervised learning for music annotation tasks. However, as supervised machine learnin…

Genre classificationInformation RetrievalMusic Genre ClassificationMusic Information Retrieval+3

An Experimental Comparison Of Multi-view Self-supervised Methods For Music Tagging

2024-04-14 · Gabriel Meseguer-Brocal, Dorian Desblancs, Romain Hennequin

Self-supervised learning has emerged as a powerful way to pre-train generalizable machine learning models on large amounts of unlabeled data. It is particularly compelling in the music domain, where obtaining labeled dat…

Contrastive LearningMusic TaggingSelf-Supervised Learning

Semi-Supervised Contrastive Learning of Musical Representations

2024-07-18 · Julien Guinot, Elio Quinton, György Fazekas

Despite the success of contrastive learning in Music Information Retrieval, the inherent ambiguity of contrastive self-supervision presents a challenge. Relying solely on augmentation chains and self-supervised positive …

Contrastive LearningInformation RetrievalMusic Information RetrievalTransfer Learning

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

2025-01-02 · Haina Zhu, Yizhi Zhou, Hangting Chen, Jianwei Yu 외

Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music tagging, instrument classification, key detec…

Contrastive LearningKey DetectionMusic TaggingQuantization+2