Class-conditional embeddings for music source separation
Isolating individual instruments in a musical mixture has a myriad of potential applications, and seems imminently achievable given the levels of performance reached by recent deep learning methods. While most musical source separation techniques learn an independent model for each instrument, we propose using a common embedding space for the time-frequency bins of all instruments in a mixture inspired by deep clustering and deep attractor networks. Additionally, an auxiliary network is used to generate parameters of a Gaussian mixture model (GMM) where the posterior distribution over GMM components in the embedding space can be used to create a mask that separates individual sources from a mixture. In addition to outperforming a mask-inference baseline on the MUSDB-18 dataset, our embedding space is easily interpretable and can be used for query-based separation.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDeep ClusteringMusic Source SeparationSimilar Papers 제목 키워드 기반
SingSong: Generating musical accompaniments from singing
We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create music featuring their own voice. To accompl…
Audio GenerationRetrievalVisual Scene Graphs for Audio Source Separation
State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sou…
Audio Source SeparationVisually Guided Sound Source SeparationImproving Universal Sound Separation Using Sound Classification
Deep learning approaches have recently achieved impressive performance on both audio source separation and sound classification. Most audio source separation approaches focus only on separating sources belonging to a res…
Audio Source SeparationClassificationGeneral ClassificationSound ClassificationPre-training Music Classification Models via Music Source Separation
In this paper, we study whether music source separation can be used as a pre-training strategy for music representation learning, targeted at music classification tasks. To this end, we first pre-train U-Net networks und…
ClassificationGenre classificationMusic Auto-TaggingMusic Classification+3MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source Extraction
We present MGE-LDM, a unified latent diffusion framework for simultaneous music generation, source imputation, and query-driven source separation. Unlike prior approaches constrained to fixed instrument classes, MGE-LDM …
ImputationMusic Generation