Why does music source separation benefit from cacophony?
In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These random mixes have mismatched characteristics compared to real music, e.g., the different stems do not have consistent beat or tonality, resulting in a cacophony. In this work, we investigate why random mixing is effective when training a state-of-the-art music source separation model in spite of the apparent distribution shift it creates. Additionally, we examine why performance levels off despite potentially limitless combinations, and examine the sensitivity of music source separation performance to differences in beat and tonality of the instrumental sources in a mixture.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationMusic Source SeparationSimilar Papers 제목 키워드 기반
Unsupervised Source Separation By Steering Pretrained Music Models
We showcase an unsupervised method that repurposes deep models trained for music generation and music tagging for audio source separation, without any retraining. An audio generation model is conditioned on an input mixt…
Audio GenerationAudio Source SeparationMusic GenerationMusic Tagging+1A Hands-on Comparison of DNNs for Dialog Separation Using Transfer Learning from Music Source Separation
This paper describes a hands-on comparison on using state-of-the-art music source separation deep neural networks (DNNs) before and after task-specific fine-tuning for separating speech content from non-speech content in…
Music Source SeparationTransfer LearningA Knowledge-Driven Approach to Music Segmentation, Music Source Separation and Cinematic Audio Source Separation
We propose a knowledge-driven, model-based approach to segmenting audio into single-category and mixed-category chunks with applications to source separation. "Knowledge" here denotes information associated with the data…
Music Source SeparationAudio Source SeparationAdversarial Semi-Supervised Audio Source Separation applied to Singing Voice Extraction
The state of the art in music source separation employs neural networks trained in a supervised fashion on multi-track databases to estimate the sources from a given mixture. With only few datasets available, often exten…
Audio Source SeparationData AugmentationMusic Source SeparationMusic Source Separation in the Waveform Domain
Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any othe…
Audio GenerationAudio SynthesisData AugmentationMulti-task Audio Source Seperation+2