paper-with-me

Papers

Interrupted and cascaded permutation invariant training for speech separation

2019-10-28 · Gene-Ping Yang, Szu-Lin Wu, Yao-Wen Mao, Hung-Yi Lee, Lin-shan Lee

Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the minimum cost label assignments dynamically, very few studies considered the separation problem to be optimizing both the model parameters and the label assignments, but focused on searching for good model architecture and parameters. In this paper, we investigate instead for a given model architecture the various flexible label assignment strategies for training the model, rather than directly using PIT. Surprisingly, we discover a significant performance boost compared to PIT is possible if the model is trained with fixed label assignments and a good set of labels is chosen. With fixed label training cascaded between two sections of PIT, we achieved the state-of-the-art performance on WSJ0-2mix without changing the model architecture at all.

📄 PDF Abstract BibTeX arXiv:1910.12706

Code (1)

r06944010/Speech-Separation-TF2 공식 구현 tf

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Recognizing Multi-talker Speech with Permutation Invariant Training

2017-03-22 · Dong Yu, Xuankai Chang, Yanmin Qian

In this paper, we propose a novel technique for direct recognition of multiple speech streams given the single channel of mixed speech, without first separating them. Our technique is based on permutation invariant train…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Many-Speakers Single Channel Speech Separation with Optimal Permutation Training

2021-04-18 · Shaked Dovrat, Eliya Nachmani, Lior Wolf

Single channel speech separation has experienced great progress in the last few years. However, training neural speech separation for a large number of speakers (e.g., more than 10 speakers) is out of reach for the curre…

Speech Separation

Single-channel speech separation using Soft-minimum Permutation Invariant Training

2021-11-16 · Midia Yousefi, John H. L. Hansen

The goal of speech separation is to extract multiple speech sources from a single microphone recording. Recently, with the advancement of deep learning and availability of large datasets, speech separation has been formu…

Speech Separation

MixCycle: Unsupervised Speech Separation via Cyclic Mixture Permutation Invariant Training

2022-02-08 · Ertuğ Karamatlı, Serap Kırbız

We introduce two unsupervised source separation methods, which involve self-supervised training from single-channel two-source speech mixtures. Our first method, mixture permutation invariant training (MixPIT), enables l…

Data AugmentationSpeech Separation

On permutation invariant training for speech source separation

2021-02-09 · Xiaoyu Liu, Jordi Pons

We study permutation invariant training (PIT), which targets at the permutation ambiguity problem for speaker independent source separation models. We extend two state-of-the-art PIT strategies. First, we look at the two…

ClusteringSpeaker Separation