RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
In this paper, we aim to generate clean speech frame by frame from a live video stream and a noisy audio stream without relying on future inputs. To this end, we propose RT-LA-VocE, which completely re-designs every component of LA-VocE, a state-of-the-art non-causal audio-visual speech enhancement model, to perform causal real-time inference with a 40ms input frame. We do so by devising new visual and audio encoders that rely solely on past frames, replacing the Transformer encoder with the Emformer, and designing a new causal neural vocoder C-HiFi-GAN. On the popular AVSpeech dataset, we show that our algorithm achieves state-of-the-art results in all real-time scenarios. More importantly, each component is carefully tuned to minimize the algorithm latency to the theoretical minimum (40ms) while maintaining a low end-to-end processing latency of 28.15ms per frame, enabling real-time frame-by-frame enhancement with minimal delay.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech EnhancementMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
LA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders
Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvement…
Speech EnhancementSpeech SynthesisNasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
Silent and whispered speech offer promise for always-available voice interaction with AI, yet existing methods struggle to balance vocabulary size, wearability, silence, and noise robustness. We present NasoVoce, a nose-…
Real-Time System for Audio-Visual Target Speech Enhancement
We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…
Audio-Visual Speech RecognitionSpeech EnhancementAV Taris: Online Audio-Visual Speech Recognition
In recent years, Automatic Speech Recognition (ASR) technology has approached human-level performance on conversational speech under relatively clean listening conditions. In more demanding situations involving distant m…
Action DetectionActivity DetectionAudio-Visual Speech RecognitionAutomatic Speech Recognition+4Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAV…
Audio-Visual Speech RecognitionActive Speaker DetectionSpeech Enhancement