On Training Targets and Objective Functions for Deep-Learning-Based Audio-Visual Speech Enhancement
Audio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been adopted to solve the AV-SE task in a supervised manner. In this context, the choice of the target, i.e. the quantity to be estimated, and the objective function, which quantifies the quality of this estimate, to be used for training is critical for the performance. This work is the first that presents an experimental study of a range of different targets and objective functions used to train a deep-learning-based AV-SE system. The results show that the approaches that directly estimate a mask perform the best overall in terms of estimated speech quality and intelligibility, although the model that directly estimates the log magnitude spectrum performs as good in terms of estimated speech quality.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningSpeech EnhancementSimilar Papers 제목 키워드 기반
Few-Shot Audio-Visual Learning of Environment Acoustics
Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional me…
audio-visual learningRoom Impulse Response (RIR)Jointly Learning Visual and Auditory Speech Representations from Raw Data
We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…
Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, th…
Deep LearningSpeech EnhancementSpeech SeparationMulti-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous rese…
Audio-Visual Speech RecognitionMulti-Task LearningRepresentation Learningspeech-recognition+3Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
Audio-visual correlation learning aims to capture essential correspondences and understand natural phenomena between audio and video. With the rapid growth of deep learning, an increasing amount of attention has been pai…
Survey