paper-with-me

Papers

Towards Intelligibility-Oriented Audio-Visual Speech Enhancement

2021-11-18 · Tassadaq Hussain, Mandar Gogate, Kia Dashtipour, Amir Hussain

Existing deep learning (DL) based speech enhancement approaches are generally optimised to minimise the distance between clean and enhanced speech features. These often result in improved speech quality however they suffer from a lack of generalisation and may not deliver the required speech intelligibility in real noisy situations. In an attempt to address these challenges, researchers have explored intelligibility-oriented (I-O) loss functions and integration of audio-visual (AV) information for more robust speech enhancement (SE). In this paper, we introduce DL based I-O SE algorithms exploiting AV information, which is a novel and previously unexplored research direction. Specifically, we present a fully convolutional AV SE model that uses a modified short-time objective intelligibility (STOI) metric as a training cost function. To the best of our knowledge, this is the first work that exploits the integration of AV modalities with an I-O based loss function for SE. Comparative experimental results demonstrate that our proposed I-O AV SE framework outperforms audio-only (AO) and AV models trained with conventional distance-based loss functions, in terms of standard objective evaluation measures when dealing with unseen speakers and noises.

📄 PDF Abstract BibTeX arXiv:2111.09642

Code (1)

cogmhear/Intelligibility-Oriented-Audio-Visual-Speech-Enhancement 공식 구현 pytorch

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Deep-Learning-Based Audio-Visual Speech Enhancement in Presence of Lombard Effect

2019-05-29 · Daniel Michelsanti, Zheng-Hua Tan, Sigurdur Sigurdsson, Jesper Jensen

When speaking in presence of background noise, humans reflexively change their way of speaking in order to improve the intelligibility of their speech. This reflex is known as Lombard effect. Collecting speech in Lombard…

Speech Enhancement

Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation

2023-05-24 · Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

Audio-visual speech enhancement (AV-SE) aims to enhance degraded speech along with extra visual information such as lip videos, and has been shown to be more effective than audio-only speech enhancement. This paper propo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge DistillationSpeech Enhancement+2

Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues

2024-02-26 · Tassadaq Hussain, Kia Dashtipour, Yu Tsao, Amir Hussain

In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often f…

DecoderSpeech Enhancement

On Training Targets and Objective Functions for Deep-Learning-Based Audio-Visual Speech Enhancement

2018-11-15 · Daniel Michelsanti, Zheng-Hua Tan, Sigurdur Sigurdsson, Jesper Jensen

Audio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been ad…

Deep LearningSpeech Enhancement

Audio-Visual Speech Enhancement Using Self-supervised Learning to Improve Speech Intelligibility in Cochlear Implant Simulations

2023-07-15 · Richard Lee Lai, Jen-Cheng Hou, I-Chun Chern, Kuo-Hsuan Hung 외

Individuals with hearing impairments face challenges in their ability to comprehend speech, particularly in noisy environments. The aim of this study is to explore the effectiveness of audio-visual speech enhancement (AV…

Self-Supervised LearningSpeech Enhancement