paper-with-me

Papers

aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio

2024-09-05 · Yan Ru Pei, Ritik Shrivastava, FNU Sidharth

We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with additional assessments on tasks such as super-resolution and de-quantization. We benchmark aTENNuate on the VoiceBank + DEMAND and the Microsoft DNS1 synthetic test sets. The network outperforms previous real-time denoising models in terms of PESQ score, parameter count, MACs, and latency. Even as a raw waveform processing model, the model maintains high fidelity to the clean signal with minimal audible artifacts. In addition, the model remains performant even when the noisy input is compressed down to 4000Hz and 4 bits, suggesting general speech enhancement capabilities in low-resource environments. Try it out by pip install attenuate

📄 PDF Abstract BibTeX arXiv:2409.03377

Code (0)

등록된 구현이 없습니다.

Tasks

Audio DenoisingDenoisingSpeech DenoisingSpeech EnhancementSuper-Resolution

Similar Papers 제목 키워드 기반

Speech enhancement aided end-to-end multi-task learning for voice activity detection

2020-10-23 · Xu Tan, Xiao-Lei Zhang

Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address…

Action DetectionActivity DetectionDecoderMulti-Task Learning+2

Attention-based Speech Enhancement Using Human Quality Perception Modelling

2023-03-23 · Khandokar Md. Nayem, Donald S. Williamson

Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize…

Language ModelingLanguage ModellingSpeech Enhancement

Trainable Adaptive Window Switching for Speech Enhancement

2018-11-05 · Yuma Koizumi, Noboru Harada, Yoichi Haneda

This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask proces…

Speech Enhancement

Real-Time Streamable Generative Speech Restoration with Flow Matching

2025-12-22 · Simon Welker, Bunlong Lay, Maris Hillemann, Tal Peer 외 arxiv

Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their application in real-time communication …

Bandwidth ExtensionSpeech Enhancement

Real Time Speech Enhancement in the Waveform Domain

2020-06-23 · Alexandre Defossez, Gabriel Synnaeve, Yossi Adi

We present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on bo…

CPUData AugmentationDecoderSpeech Enhancement