Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
We introduce a new paradigm for active sound modification: Active Speech Enhancement (ASE). While Active Noise Cancellation (ANC) algorithms focus on suppressing external interference, ASE goes further by actively shaping the speech signal -- both attenuating unwanted noise components and amplifying speech-relevant frequencies -- to improve intelligibility and perceptual quality. To enable this, we propose a novel Transformer-Mamba-based architecture, along with a task-specific loss function designed to jointly optimize interference suppression and signal enrichment. Our method outperforms existing baselines across multiple speech processing tasks -- including denoising, dereverberation, and declipping -- demonstrating the effectiveness of active, targeted modulation in challenging acoustic environments.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingMambaSpeech DenoisingSpeech EnhancementMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning
In this paper, we explore a continuous modeling approach for deep-learning-based speech enhancement, focusing on the denoising process. We use a state variable to indicate the denoising process. The starting state is noi…
Automatic Speech RecognitionDenoisingSpeech Enhancementspeech-recognition+1Look\&Listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement
Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed a…
Active Speaker DetectionMulti-Task LearningSpeech EnhancementZero-Shot Personalized Speech Enhancement through Speaker-Informed Model Selection
This paper presents a novel zero-shot learning approach towards personalized speech enhancement through the use of a sparsely active ensemble model. Optimizing speech denoising systems towards a particular test-time spea…
ClusteringDenoisingModel SelectionSpeech Denoising+2Real-Time System for Audio-Visual Target Speech Enhancement
We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…
Audio-Visual Speech RecognitionSpeech EnhancementAeGAN: Time-Frequency Speech Denoising via Generative Adversarial Networks
Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in rea…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingGenerative Adversarial Network+6