paper-with-me

홈 › Papers

Text-aware Speech Separation for Multi-talker Keyword Spotting

2024-06-18 · Haoyu Li, Baochen Yang, Yu Xi, Linfeng Yu, Tian Tan, Hao Li, Kai Yu

For noisy environments, ensuring the robustness of keyword spotting (KWS) systems is essential. While much research has focused on noisy KWS, less attention has been paid to multi-talker mixed speech scenarios. Unlike the usual cocktail party problem where multi-talker speech is separated using speaker clues, the key challenge here is to extract the target speech for KWS based on text clues. To address it, this paper proposes a novel Text-aware Permutation Determinization Training method for multi-talker KWS with a clue-based Speech Separation front-end (TPDT-SS). Our research highlights the critical role of SS front-ends and shows that incorporating keyword-specific clues into these models can greatly enhance the effectiveness. TPDT-SS shows remarkable success in addressing permutation problems in mixed keyword speech, thereby greatly boosting the performance of the backend. Additionally, fine-tuning our system on unseen mixed speech results in further performance improvement.

📄 PDF Abstract BibTeX arXiv:2406.12447

Code (1)

gnafiy/tpdt-ss-kws 공식 구현 pytorch

Tasks

Keyword SpottingSpeech Separation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Multi-talker ASR for an unknown number of sources: Joint training of source counting, separation and ASR

2020-06-04 · Thilo von Neumann, Christoph Boeddeker, Lukas Drude, Keisuke Kinoshita 외

Most approaches to multi-talker overlapped speech separation and recognition assume that the number of simultaneously active speakers is given, but in realistic situations, it is typically unknown. To cope with this, we …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Extractionspeech-recognition+2

Utterance-level Permutation Invariant Training with Latency-controlled BLSTM for Single-channel Multi-talker Speech Separation

2019-12-25 · Lu Huang, Gaofeng Cheng, Pengyuan Zhang, Yi Yang 외

Utterance-level permutation invariant training (uPIT) has achieved promising progress on single-channel multi-talker speech separation task. Long short-term memory (LSTM) and bidirectional LSTM (BLSTM) are widely used as…

Speech Separation

Single-Channel Multi-talker Speech Recognition with Permutation Invariant Training

2017-07-19 · Yanmin Qian, Xuankai Chang, Dong Yu

Although great progresses have been made in automatic speech recognition (ASR), significant performance degradation is still observed when recognizing multi-talker mixed speech. In this paper, we propose and evaluate sev…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Multi-channel Conversational Speaker Separation via Neural Diarization

2023-11-15 · Hassan Taherian, DeLiang Wang

When dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or mee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognition+1

Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording

2020-12-17 · Cong Han, Yi Luo, Chenda Li, Tianyan Zhou 외

Leveraging additional speaker information to facilitate speech separation has received increasing attention in recent years. Recent research includes extracting target speech by using the target speaker's voice snippet a…

ClusteringSpeech Separation