paper-with-me

Papers

Unsupervised Sound Separation Using Mixture Invariant Training

2020-06-23 · NeurIPS 2020 12 · Scott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss, Kevin Wilson, John R. Hershey

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from synthetic mixtures created by adding up isolated ground-truth sources. Reliance on this synthetic training data is problematic because good performance depends upon the degree of match between the training data and real-world audio, especially in terms of the acoustic conditions and distribution of sources. The acoustic properties can be challenging to accurately simulate, and the distribution of sound types may be hard to replicate. In this paper, we propose a completely unsupervised method, mixture invariant training (MixIT), that requires only single-channel acoustic mixtures. In MixIT, training examples are constructed by mixing together existing mixtures, and the model separates them into a variable number of latent sources, such that the separated sources can be remixed to approximate the original mixtures. We show that MixIT can achieve competitive performance compared to supervised methods on speech separation. Using MixIT in a semi-supervised learning setting enables unsupervised domain adaptation and learning from large amounts of real world data without ground-truth source waveforms. In particular, we significantly improve reverberant speech separation performance by incorporating reverberant mixtures, train a speech enhancement system from noisy mixtures, and improve universal sound separation by incorporating a large amount of in-the-wild data.

📄 PDF Abstract BibTeX arXiv:2006.12701

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationSpeech EnhancementSpeech SeparationUnsupervised Domain Adaptation

Similar Papers 제목 키워드 기반

Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

2020-11-02 · ICLR 2021 1 · Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey 외

Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural videos remains an open problem. In this work,…

Scene Understanding

Sparse, Efficient, and Semantic Mixture Invariant Training: Taming In-the-Wild Unsupervised Sound Separation

2021-06-01 · Scott Wisdom, Aren Jansen, Ron J. Weiss, Hakan Erdogan 외

Supervised neural network training has led to significant progress on single-channel sound separation. This approach relies on ground truth isolated sources, which precludes scaling to widely available mixture data and l…

MixCycle: Unsupervised Speech Separation via Cyclic Mixture Permutation Invariant Training

2022-02-08 · Ertuğ Karamatlı, Serap Kırbız

We introduce two unsupervised source separation methods, which involve self-supervised training from single-channel two-source speech mixtures. Our first method, mixture permutation invariant training (MixPIT), enables l…

Data AugmentationSpeech Separation

Improving Bird Classification with Unsupervised Sound Separation

2021-10-07 · Tom Denton, Scott Wisdom, John R. Hershey

This paper addresses the problem of species classification in bird song recordings. The massive amount of available field recordings of birds presents an opportunity to use machine learning to automatically track bird po…

Classification

Self-Supervised Learning from Automatically Separated Sound Scenes

2021-05-05 · Eduardo Fonseca, Aren Jansen, Daniel P. W. Ellis, Scott Wisdom 외

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events wit…

Contrastive LearningSelf-Supervised Learning