paper-with-me

Papers

Guided Training: A Simple Method for Single-channel Speaker Separation

2021-03-26 · Hao Li, Xueliang Zhang, Guanglai Gao

Deep learning has shown a great potential for speech separation, especially for speech and non-speech separation. However, it encounters permutation problem for multi-speaker separation where both target and interference are speech. Permutation Invariant training (PIT) was proposed to solve this problem by permuting the order of the multiple speakers. Another way is to use an anchor speech, a short speech of the target speaker, to model the speaker identity. In this paper, we propose a simple strategy to train a long short-term memory (LSTM) model to solve the permutation problem in speaker separation. Specifically, we insert a short speech of target speaker at the beginning of a mixture as guide information. So, the first appearing speaker is defined as the target. Due to the powerful capability on sequence modeling, LSTM can use its memory cells to track and separate target speech from interfering speech. Experimental results show that the proposed training strategy is effective for speaker separation.

📄 PDF Abstract BibTeX arXiv:2103.14330

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker SeparationSpeech Separation

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Beamformer-Guided Target Speaker Extraction

2023-03-15 · Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets

We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs …

Target Speaker Extraction

End-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis

2023-10-16 · Can Cui, Imran Ahamad Sheikh, Mostafa Sadeghi, Emmanuel Vincent

We present an end-to-end multichannel speaker-attributed automatic speech recognition (MC-SA-ASR) system that combines a Conformer-based encoder with multi-frame crosschannel attention and a speaker-attributed Transforme…

Automatic Speech RecognitionDecoderSpeaker Identificationspeech-recognition+1

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

2023-11-01 · Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi 외

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech-to-Text+2

Mutual Learning of Single- and Multi-Channel End-to-End Neural Diarization

2022-10-07 · Shota Horiguchi, Yuki Takashima, Shinji Watanabe, Paola Garcia

Due to the high performance of multi-channel speech processing, we can use the outputs from a multi-channel model as teacher labels when training a single-channel model with knowledge distillation. To the contrary, it is…

Knowledge Distillationspeaker-diarizationSpeaker DiarizationTransfer Learning

Multi-Channel Speaker Verification for Single and Multi-talker Speech

2020-10-23 · Saurabh Kataria, Shi-Xiong Zhang, Dong Yu

To improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features. Specifically, we combine spectral, …

Action DetectionActivity DetectionSpeaker VerificationSpeech Enhancement+1