paper-with-me

Papers

Multi-channel Speech Separation Using Deep Embedding Model with Multilayer Bootstrap Networks

2019-10-24 · Ziye Yang, Xiao-Lei Zhang

Recently, deep clustering (DPCL) based speaker-independent speech separation has drawn much attention, since it needs little speaker prior information. However, it still has much room of improvement, particularly in reverberant environments. If the training and test environments mismatch which is a common case, the embedding vectors produced by DPCL may contain much noise and many small variations. To deal with the problem, we propose a variant of DPCL, named DPCL++, by applying a recent unsupervised deep learning method---multilayer bootstrap networks(MBN)---to further reduce the noise and small variations of the embedding vectors in an unsupervised way in the test stage, which fascinates k-means to produce a good result. MBN builds a gradually narrowed network from bottom-up via a stack of k-centroids clustering ensembles, where the k-centroids clusterings are trained independently by random sampling and one-nearest-neighbor optimization. To further improve the robustness of DPCL++ in reverberant environments, we take spatial features as part of its input. Experimental results demonstrate the effectiveness of the proposed method.

📄 PDF Abstract BibTeX arXiv:1910.10912

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDeep ClusteringSpeech Separation

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

CasNet: Investigating Channel Robustness for Speech Separation

2022-10-27 · Fan-Lin Wang, Yao-Fei Cheng, Hung-Shin Lee, Yu Tsao 외

Recording channel mismatch between training and testing conditions has been shown to be a serious problem for speech separation. This situation greatly reduces the separation performance, and cannot meet the requirement …

Speech Separation

Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism

2021-02-07 · Jisi Zhang, Catalin Zorila, Rama Doddipatla, Jon Barker

In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on a…

Speech Extractionspeech-recognitionSpeech RecognitionSpeech Separation

MIMO-DBnet: Multi-channel Input and Multiple Outputs DOA-aware Beamforming Network for Speech Separation

2022-12-07 · Yanjie Fu, Haoran Yin, Meng Ge, Longbiao Wang 외

Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speaker feature, face image or directional in…

Speech Separation

Single-Channel Speech Separation with Auxiliary Speaker Embeddings

2019-06-24 · Shuo Liu, Gil Keren, Björn Schuller

We present a novel source separation model to decompose asingle-channel speech signal into two speech segments belonging to two different speakers. The proposed model is a neural network based on residual blocks, and use…

Speech Separation

Royalflush Speaker Diarization System for ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge

2022-02-10 · Jingguang Tian, Xinhui Hu, Xinkang Xu

This paper describes the Royalflush speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription Challenge(M2MeT). Our system comprises speech enhancement, overlapped speech detection, spea…

speaker-diarizationSpeaker DiarizationSpeaker VerificationSpeech Enhancement+1