paper-with-me

홈 › Papers

Content based singing voice source separation via strong conditioning using aligned phonemes

2020-08-05 · Gabriel Meseguer-Brocal, Geoffroy Peeters

Informed source separation has recently gained renewed interest with the introduction of neural networks and the availability of large multitrack datasets containing both the mixture and the separated sources. These approaches use prior information about the target source to improve separation. Historically, Music Information Retrieval researchers have focused primarily on score-informed source separation, but more recent approaches explore lyrics-informed source separation. However, because of the lack of multitrack datasets with time-aligned lyrics, models use weak conditioning with non-aligned lyrics. In this paper, we present a multimodal multitrack dataset with lyrics aligned in time at the word level with phonetic information as well as explore strong conditioning using the aligned phonemes. Our model follows a U-Net architecture and takes as input both the magnitude spectrogram of a musical mixture and a matrix with aligned phonetic information. The phoneme matrix is embedded to obtain the parameters that control Feature-wise Linear Modulation (FiLM) layers. These layers condition the U-Net feature maps to adapt the separation process to the presence of different phonemes via affine transformations. We show that phoneme conditioning can be successfully applied to improve singing voice source separation.

📄 PDF Abstract BibTeX arXiv:2008.02070

Code (1)

gabolsgabs/vunet 공식 구현 tf

Tasks

Information RetrievalMusic Information RetrievalRetrieval

Methods 이 논문이 사용한 방법론

Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

MedleyVox: An Evaluation Dataset for Multiple Singing Voices Separation

2022-11-14 · Chang-Bin Jeon, Hyeongi Moon, Keunwoo Choi, Ben Sangbae Chon 외

Separation of multiple singing voices into each voice is a rarely studied area in music source separation research. The absence of a benchmark dataset has hindered its progress. In this paper, we present an evaluation da…

Music Source SeparationSuper-Resolution

A cappella: Audio-visual Singing Voice Separation

2021-04-20 · Juan F. Montesinos, Venkatesh S. Kadandale, Gloria Haro

The task of isolating a target singing voice in music videos has useful applications. In this work, we explore the single-channel singing voice separation problem from a multimodal perspective, by jointly learning from a…

Music Source SeparationSpeech Separation

Towards Reliable Objective Evaluation Metrics for Generative Singing Voice Separation Models

2025-07-15 · Paul A. Bereuter, Benjamin Stahl, Mark D. Plumbley, Alois Sontacchi

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative mod…

Audio Source Separationblind source separation

Unsupervised Interpretable Representation Learning for Singing Voice Separation

2020-03-03 · Stylianos I. Mimilakis, Konstantinos Drossos, Gerald Schuller

In this work, we present a method for learning interpretable music signal representations directly from waveform signals. Our method can be trained using unsupervised objectives and relies on the denoising auto-encoder m…

DenoisingMusic Source SeparationRepresentation Learning

Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation

2015-02-13 · Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, Paris Smaragdis

Monaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possi…

DenoisingSpeech DenoisingSpeech Separation