paper-with-me

Papers

Multi-Channel Target Speaker Extraction with Refinement: The WavLab Submission to the Second Clarity Enhancement Challenge

2023-02-15 · Samuele Cornell, Zhong-Qiu Wang, Yoshiki Masuyama, Shinji Watanabe, Manuel Pariente, Nobutaka Ono

This paper describes our submission to the Second Clarity Enhancement Challenge (CEC2), which consists of target speech enhancement for hearing-aid (HA) devices in noisy-reverberant environments with multiple interferers such as music and competing speakers. Our approach builds upon the powerful iterative neural/beamforming enhancement (iNeuBe) framework introduced in our recent work, and this paper extends it for target speaker extraction. We therefore name the proposed approach as iNeuBe-X, where the X stands for extraction. To address the challenges encountered in the CEC2 setting, we introduce four major novelties: (1) we extend the state-of-the-art TF-GridNet model, originally designed for monaural speaker separation, for multi-channel, causal speech enhancement, and large improvements are observed by replacing the TCNDenseNet used in iNeuBe with this new architecture; (2) we leverage a recent dual window size approach with future-frame prediction to ensure that iNueBe-X satisfies the 5 ms constraint on algorithmic latency required by CEC2; (3) we introduce a novel speaker-conditioning branch for TF-GridNet to achieve target speaker extraction; (4) we propose a fine-tuning step, where we compute an additional loss with respect to the target speaker signal compensated with the listener audiogram. Without using external data, on the official development set our best model reaches a hearing-aid speech perception index (HASPI) score of 0.942 and a scale-invariant signal-to-distortion ratio improvement (SI-SDRi) of 18.8 dB. These results are promising given the fact that the CEC2 data is extremely challenging (e.g., on the development set the mixture SI-SDR is -12.3 dB). A demo of our submitted system is available at WAVLab CEC2 demo.

📄 PDF Abstract BibTeX arXiv:2302.07928

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker SeparationSpeech EnhancementTarget Speaker Extraction

Similar Papers 제목 키워드 기반

Beamformer-Guided Target Speaker Extraction

2023-03-15 · Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets

We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs …

Target Speaker Extraction

Deep Ad-hoc Beamforming Based on Speaker Extraction for Target-Dependent Speech Separation

2020-12-01 · Ziye Yang, Shanzheng Guan, Xiao-Lei Zhang

Recently, the research on ad-hoc microphone arrays with deep learning has drawn much attention, especially in speech enhancement and separation. Because an ad-hoc microphone array may cover such a large area that multipl…

channel selectionDeep LearningSpeech EnhancementSpeech Separation

Binaural Selective Attention Model for Target Speaker Extraction

2024-06-18 · Hanyu Meng, Qiquan Zhang, Xiangyu Zhang, Vidhyasaharan Sethu 외

The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extractio…

modelTarget Speaker Extraction

L-SpEx: Localized Target Speaker Extraction

2022-02-21 · Meng Ge, Chenglin Xu, Longbiao Wang, Eng Siong Chng 외

Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction…

Target Speaker Extraction

Speaker-conditioned Target Speaker Extraction based on Customized LSTM Cells

2021-04-09 · Ragini Sinha, Marvin Tammen, Christian Rollwage, Simon Doclo

Speaker-conditioned target speaker extraction systems rely on auxiliary information about the target speaker to extract the target speaker signal from a mixture of multiple speakers. Typically, a deep neural network is a…

Target Speaker Extraction