paper-with-me

Papers

MixCycle: Unsupervised Speech Separation via Cyclic Mixture Permutation Invariant Training

2022-02-08 · Ertuğ Karamatlı, Serap Kırbız

We introduce two unsupervised source separation methods, which involve self-supervised training from single-channel two-source speech mixtures. Our first method, mixture permutation invariant training (MixPIT), enables learning a neural network model which separates the underlying sources via a challenging proxy task without supervision from the reference sources. Our second method, cyclic mixture permutation invariant training (MixCycle), uses MixPIT as a building block in a cyclic fashion for continuous learning. MixCycle gradually converts the problem from separating mixtures of mixtures into separating single mixtures. We compare our methods to common supervised and unsupervised baselines: permutation invariant training with dynamic mixing (PIT-DM) and mixture invariant training (MixIT). We show that MixCycle outperforms MixIT and reaches a performance level very close to the supervised baseline (PIT-DM) while circumventing the over-separation issue of MixIT. Also, we propose a self-evaluation technique inspired by MixCycle that estimates model performance without utilizing any reference sources. We show that it yields results consistent with an evaluation on reference sources (LibriMix) and also with an informal listening test conducted on a real-life mixtures dataset (REAL-M).

📄 PDF Abstract BibTeX arXiv:2202.03875

Code (1)

ertug/mixcycle 공식 구현 pytorch

Tasks

Data AugmentationSpeech Separation

Similar Papers 제목 키워드 기반

Self-Remixing: Unsupervised Speech Separation via Separation and Remixing

2022-11-18 · Kohei Saijo, Tetsuji Ogawa

We present Self-Remixing, a novel self-supervised speech separation method, which refines a pre-trained separation model in an unsupervised manner. The proposed method consists of a shuffler module and a solver module, a…

Domain AdaptationSemi-supervised Domain AdaptationSpeech Separation

UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training Mixtures

2023-05-31 · NeurIPS 2023 11

In reverberant conditions with multiple concurrent speakers, each microphone acquires a mixture signal of multiple speakers at a different location. In over-determined conditions where the microphones out-number speakers…

Speaker SeparationSpeech Separation

Unsupervised Sound Separation Using Mixture Invariant Training

2020-06-23 · NeurIPS 2020 12 · Scott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss 외

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the componen…

Domain AdaptationSpeech EnhancementSpeech SeparationUnsupervised Domain Adaptation

Heterogeneous Separation Consistency Training for Adaptation of Unsupervised Speech Separation

2022-04-23 · Jiangyu Han, Yanhua Long

Recently, supervised speech separation has made great progress. However, limited by the nature of supervised training, most existing separation methods require ground-truth sources and are trained on synthetic datasets. …

Speech Separation

Enhanced Reverberation as Supervision for Unsupervised Speech Separation

2024-08-06 · Kohei Saijo, Gordon Wichern, François G. Germain, Zexu Pan 외

Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models are trained so that sources predicted fro…

Speech Separation