paper-with-me

Papers

Audio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras

2019-12-05 · Ander Arriandiaga, Giovanni Morrone, Luca Pasa, Leonardo Badino, Chiara Bartolozzi

We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is the movement of the facial landmarks related to speech production. However, all approaches proposed so far work offline, using frame-based video input, making it difficult to process an audio-visual signal with low latency, for online applications. In order to overcome this limitation, we propose the use of event-driven cameras and exploit compression, high temporal resolution and low latency, for low cost and low latency motion feature extraction, going towards online embedded audio-visual speech processing. We use the event-driven optical flow estimation of the facial landmarks as input to a stacked Bidirectional LSTM trained to predict an Ideal Amplitude Mask that is then used to filter the noisy audio, to obtain the audio signal of the target speaker. The presented approach performs almost on par with the frame-based approach, with very low latency and computational cost.

📄 PDF Abstract BibTeX arXiv:1912.02671

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationSpeech Separation

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

2025-07-29 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAV…

Audio-Visual Speech RecognitionActive Speaker DetectionSpeech Enhancement

Look\&Listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement

2022-03-04 · Junwen Xiong, Yu Zhou, Peng Zhang, Lei Xie 외

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed a…

Active Speaker DetectionMulti-Task LearningSpeech Enhancement

Real-Time System for Audio-Visual Target Speech Enhancement

2025-09-25 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…

Audio-Visual Speech RecognitionSpeech Enhancement

LA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders

2022-11-20 · Rodrigo Mira, Buye Xu, Jacob Donley, Anurag Kumar 외

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvement…

Speech EnhancementSpeech Synthesis

The Conversation: Deep Audio-Visual Speech Enhancement

2018-04-11 · Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In th…

Speech Enhancement