paper-with-me

Papers

Efficient Trainable Front-Ends for Neural Speech Enhancement

2020-02-20 · Jonah Casebeer, Umut Isik, Shrikant Venkataramani, Arvindh Krishnaswamy

Many neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are implemented as large Discrete Fourier Transform matrices; which are prohibitively inefficient for low-compute systems. We present an efficient, trainable front-end based on the butterfly mechanism to compute the Fast Fourier Transform, and show its accuracy and efficiency benefits for low-compute neural speech enhancement models. We also explore the effects of making the STFT window trainable.

📄 PDF Abstract BibTeX arXiv:2002.09286

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Frontend Token Enhancement for Token-Based Speech Recognition

2026-02-04 · Takanori Ashihara, Shota Horiguchi, Kohei Matsuura, Tsubasa Ochiai 외 arxiv

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these rep…

Self-Supervised LearningSpeech Recognition

ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding

2022-07-19 · Yen-Ju Lu, Xuankai Chang, Chenda Li, Wangyou Zhang 외

This paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit. Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech RecognitionSpeech Enhancement+4

Deep Xi as a Front-End for Robust Automatic Speech Recognition

2020-01-28

Current front-ends for robust automatic speech recognition(ASR) include masking- and mapping-based deep learning approaches to speech enhancement. A recently proposed deep learning approach toa prioriSNR estimation, call…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep LearningSpeech Enhancement+2

ROBUST SPEECH COMMAND RECOGNITION USING LABEL-DRIVEN TIME-FREQUENCY MASKING

2018-10-22 · Anonymous

Speech enhancement driven robust Automatic Speech Recognition (ASR) systems typically require parallel corpus with noisy and clean speech utterances for training. Moreover, many studies have reported that such front-ends…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationSpeech Enhancement+2

Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends

2026-08-31 · Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne arxiv

Speech enhancement is often used as a front-end for robust ASR, yet recurrent temporal and cross-band modules introduce sequential dependencies that reduce parallel efficiency. In this paper, we present a sequence-parall…

Speech Enhancement