paper-with-me

Papers

DNN driven Speaker Independent Audio-Visual Mask Estimation for Speech Separation

2018-07-31 · Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker, Amir Hussain

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on target speaker while filtering out other noises. In this study, we propose a novel deep neural network (DNN) based audiovisual (AV) mask estimation model. The proposed AV mask estimation model contextually integrates the temporal dynamics of both audio and noise-immune visual features for improved mask estimation and speech separation. For optimal AV features extraction and ideal binary mask (IBM) estimation, a hybrid DNN architecture is exploited to leverages the complementary strengths of a stacked long short term memory (LSTM) and convolution LSTM network. The comparative simulation results in terms of speech quality and intelligibility demonstrate significant performance improvement of our proposed AV mask estimation model as compared to audio-only and visual-only mask estimation approaches for both speaker dependent and independent scenarios.

📄 PDF Abstract BibTeX arXiv:1808.00060

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Face Landmark-based Speaker-Independent Audio-Visual Speech Enhancement in Multi-Talker Environments

2018-11-06 · Giovanni Morrone, Luca Pasa, Vadim Tikhanoff, Sonia Bergamaschi 외

In this paper, we address the problem of enhancing the speech of a speaker of interest in a cocktail party scenario when visual information of the speaker of interest is available. Contrary to most previous studies, we d…

Speech EnhancementSpeech Separation

Audio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras

2019-12-05 · Ander Arriandiaga, Giovanni Morrone, Luca Pasa, Leonardo Badino 외

We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is th…

Optical Flow EstimationSpeech Separation

Speaker-Independent Speech-Driven Visual Speech Synthesis using Domain-Adapted Acoustic Models

2019-05-15 · Ahmed Hussen Abdelaziz, Barry-John Theobald, Justin Binder, Gabriele Fanelli 외

Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Face Modelspeech-recognition+2

SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification

2025-06-21 · Gnana Praveen Rajasekhar, Jahangir Alam

Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address th…

Contrastive LearningSelf-Supervised LearningSpeaker Verification

Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

2018-04-10 · Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel 외

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and do…

Speech Separation