paper-with-me

Papers

On Training Targets and Objective Functions for Deep-Learning-Based Audio-Visual Speech Enhancement

2018-11-15 · Daniel Michelsanti, Zheng-Hua Tan, Sigurdur Sigurdsson, Jesper Jensen

Audio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been adopted to solve the AV-SE task in a supervised manner. In this context, the choice of the target, i.e. the quantity to be estimated, and the objective function, which quantifies the quality of this estimate, to be used for training is critical for the performance. This work is the first that presents an experimental study of a range of different targets and objective functions used to train a deep-learning-based AV-SE system. The results show that the approaches that directly estimate a mask perform the best overall in terms of estimated speech quality and intelligibility, although the model that directly estimates the log magnitude spectrum performs as good in terms of estimated speech quality.

📄 PDF Abstract BibTeX arXiv:1811.06234

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningSpeech Enhancement

Similar Papers 제목 키워드 기반

Few-Shot Audio-Visual Learning of Environment Acoustics

2022-06-08 · Sagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen Grauman

Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional me…

audio-visual learningRoom Impulse Response (RIR)

Jointly Learning Visual and Auditory Speech Representations from Raw Data

2022-12-12 · Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis 외

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1

An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation

2020-08-21 · Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu 외

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, th…

Deep LearningSpeech EnhancementSpeech Separation

Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation

2025-01-23 · Sungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang 외

Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous rese…

Audio-Visual Speech RecognitionMulti-Task LearningRepresentation Learningspeech-recognition+3

Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

2022-02-28 · Luís Vilaça, Yi Yu, Paula Viana

Audio-visual correlation learning aims to capture essential correspondences and understand natural phenomena between audio and video. With the rapid growth of deep learning, an increasing amount of attention has been pai…

Survey