paper-with-me

Papers

Cleanformer: A multichannel array configuration-invariant neural enhancement frontend for ASR in smart speakers

2022-04-25 · Joseph Caroselli, Arun Narayanan, Nathan Howard, Tom O'Malley

This work introduces the Cleanformer, a streaming multichannel neural based enhancement frontend for automatic speech recognition (ASR). This model has a conformer-based architecture which takes as inputs a single channel each of raw and enhanced signals, and uses self-attention to derive a time-frequency mask. The enhanced input is generated by a multichannel adaptive noise cancellation algorithm known as Speech Cleaner, which makes use of noise context to derive its filter taps. The time-frequency mask is applied to the noisy input to produce enhanced output features for ASR. Detailed evaluations are presented with simulated and re-recorded datasets in speech-based and non-speech-based noise that show significant reduction in word error rate (WER) when using a large-scale state-of-the-art ASR model. It also will be shown to significantly outperform enhancement using a beamformer with ideal steering. The enhancement model is agnostic of the number of microphones and array configuration and, therefore, can be used with different microphone arrays without the need for retraining. It is demonstrated that performance improves with more microphones, up to 4, with each additional microphone providing a smaller marginal benefit. Specifically, for an SNR of -6dB, relative WER improvements of about 80\% are shown in both noise conditions.

📄 PDF Abstract BibTeX arXiv:2204.11933

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Spatial-Filter-Bank-Based Neural Method for Multichannel Speech Enhancement

2025-04-02 · Tianqin Zheng, Jilu Jin, Hanchen Pei, Gongping Huang 외

The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically inv…

Speech Enhancement

Array Configuration-Agnostic Personalized Speech Enhancement using Long-Short-Term Spatial Coherence

2022-11-16 · Yicheng Hsu, Yonghan Lee, Mingsian R. Bai

Personalized speech enhancement has been a field of active research for suppression of speechlike interferers such as competing speakers or TV dialogues. Compared with single channel approaches, multichannel PSE systems …

Speech Enhancement

Flexible Multichannel Speech Enhancement for Noise-Robust Frontend

2024-06-06 · Ante Jukić, Jagadeesh Balam, Boris Ginsburg

This paper proposes a flexible multichannel speech enhancement system with the main goal of improving robustness of automatic speech recognition (ASR) in noisy conditions. The proposed system combines a flexible neural m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

Advances in Microphone Array Processing and Multichannel Speech Enhancement

2025-02-13 · Gongping Huang, Jesper R. Jensen, Jingdong Chen, Jacob Benesty 외

This paper reviews pioneering works in microphone array processing and multichannel speech enhancement, highlighting historical achievements, technological evolution, commercialization aspects, and key challenges. It pro…

Speech Enhancement

Microphone Array Generalization for Multichannel Narrowband Deep Speech Enhancement

2021-07-27 · Siyuan Zhang, Xiaofei Li

This paper addresses the problem of microphone array generalization for deep-learning-based end-to-end multichannel speech enhancement. We aim to train a unique deep neural network (DNN) potentially performing well on un…

Speech Enhancement