Audio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras
We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is the movement of the facial landmarks related to speech production. However, all approaches proposed so far work offline, using frame-based video input, making it difficult to process an audio-visual signal with low latency, for online applications. In order to overcome this limitation, we propose the use of event-driven cameras and exploit compression, high temporal resolution and low latency, for low cost and low latency motion feature extraction, going towards online embedded audio-visual speech processing. We use the event-driven optical flow estimation of the facial landmarks as input to a stacked Bidirectional LSTM trained to predict an Ideal Amplitude Mask that is then used to filter the noisy audio, to obtain the audio signal of the target speaker. The presented approach performs almost on par with the frame-based approach, with very low latency and computational cost.
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Flow EstimationSpeech SeparationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAV…
Audio-Visual Speech RecognitionActive Speaker DetectionSpeech EnhancementLook\&Listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement
Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed a…
Active Speaker DetectionMulti-Task LearningSpeech EnhancementReal-Time System for Audio-Visual Target Speech Enhancement
We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…
Audio-Visual Speech RecognitionSpeech EnhancementLA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders
Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvement…
Speech EnhancementSpeech SynthesisThe Conversation: Deep Audio-Visual Speech Enhancement
Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In th…
Speech Enhancement