paper-with-me

홈 › Papers

Employing low-pass filtered temporal speech features for the training of ideal ratio mask in speech enhancement

2021-10-01 · ROCLING 2021 10 · Yan-Tong Chen, Zi-Qiang Lin, Jeih-weih Hung

The masking-based speech enhancement method pursues a multiplicative mask that applies to the spectrogram of input noise-corrupted utterance, and a deep neural network (DNN) is often used to learn the mask. In particular, the features commonly used for automatic speech recognition can serve as the input of the DNN to learn the well-behaved mask that significantly reduce the noise distortion of processed utterances. This study proposes to preprocess the input speech features for the ideal ratio mask (IRM)-based DNN by lowpass filtering in order to alleviate the noise components. In particular, we employ the discrete wavelet transform (DWT) to decompose the temporal speech feature sequence and scale down the detail coefficients, which correspond to the high-pass portion of the sequence. Preliminary experiments conducted on a subset of TIMIT corpus reveal that the proposed method can make the resulting IRM achieve higher speech quality and intelligibility for the babble noise-corrupted signals compared with the original IRM, indicating that the lowpass filtered temporal feature sequence can learn a superior IRM network for speech enhancement.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

使用低通時序列語音特徵訓練理想比率遮罩法之語音強化 (Employing Low-Pass Filtered Temporal Speech Features for the Training of Ideal Ratio Mask in Speech Enhancement)

2021-12-01 · IJCLCLP 2021 12 · Yan-Tong Chen, Jeih-weih Hung
Speech Enhancement

Adaptive Feature Selection for End-to-End Speech Translation

2020-10-16 · Findings of the Association for Computational Linguistics 2020 · Biao Zhang, Ivan Titov, Barry Haddow, Rico Sennrich

Information in speech signals is not evenly distributed, making it an additional challenge for end-to-end (E2E) speech translation (ST) to learn to focus on informative features. In this paper, we propose adaptive featur…

Data AugmentationDecoderfeature selectionTranslation

Time-Contrastive Learning Based DNN Bottleneck Features for Text-Dependent Speaker Verification

2017-04-06 · Achintya Kr. Sarkar, Zheng-Hua Tan

In this paper, we present a time-contrastive learning (TCL) based bottleneck (BN)feature extraction method for speech signals with an application to text-dependent (TD) speaker verification (SV). It is well-known that sp…

Contrastive LearningSpeaker VerificationText-Dependent Speaker Verification

Filtered Noise Shaping for Time Domain Room Impulse Response Estimation From Reverberant Speech

2021-07-15 · Christian J. Steinmetz, Vamsi Krishna Ithapu, Paul Calamia

Deep learning approaches have emerged that aim to transform an audio signal so that it sounds as if it was recorded in the same room as a reference recording, with applications both in audio post-production and augmented…

DecoderRoom Impulse Response (RIR)

Analysis of EEG frequency bands for Envisioned Speech Recognition

2022-03-29 · Ayush Tripathi

The use of Automatic speech recognition (ASR) interfaces have become increasingly popular in daily life for use in interaction and control of electronic devices. The interfaces currently being used are not feasible for a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)EEGElectroencephalogram (EEG)+2