Unsupervised Interpretable Representation Learning for Singing Voice Separation
In this work, we present a method for learning interpretable music signal representations directly from waveform signals. Our method can be trained using unsupervised objectives and relies on the denoising auto-encoder model that uses a simple sinusoidal model as decoding functions to reconstruct the singing voice. To demonstrate the benefits of our method, we employ the obtained representations to the task of informed singing voice separation via binary masking, and measure the obtained separation quality by means of scale-invariant signal to distortion ratio. Our findings suggest that our method is capable of learning meaningful representations for singing voice separation, while preserving conveniences of the the short-time Fourier transform like non-negativity, smoothness, and reconstruction subject to time-frequency masking, that are desired in audio and music source separation.
Code (1)
Tasks
DenoisingMusic Source SeparationRepresentation LearningSimilar Papers 제목 키워드 기반
A fully differentiable model for unsupervised singing voice separation
A novel model was recently proposed by Schulze-Forster et al. in [1] for unsupervised music source separation. This model allows to tackle some of the major shortcomings of existing source separation frameworks. Specific…
Music Source SeparationMedleyVox: An Evaluation Dataset for Multiple Singing Voices Separation
Separation of multiple singing voices into each voice is a rarely studied area in music source separation research. The absence of a benchmark dataset has hindered its progress. In this paper, we present an evaluation da…
Music Source SeparationSuper-ResolutionInformed Group-Sparse Representation for Singing Voice Separation
Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that t…
Information RetrievalMusic Information RetrievalRetrievalA cappella: Audio-visual Singing Voice Separation
The task of isolating a target singing voice in music videos has useful applications. In this work, we explore the single-channel singing voice separation problem from a multimodal perspective, by jointly learning from a…
Music Source SeparationSpeech SeparationPitchNet: Unsupervised Singing Voice Conversion with Pitch Adversarial Network
Singing voice conversion is to convert a singer's voice to another one's voice without changing singing content. Recent work shows that unsupervised singing voice conversion can be achieved with an autoencoder-based appr…
DecoderMusic GenerationTranslationVoice Conversion