Vocoder-Based Speech Synthesis from Silent Videos
Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way to synthesise speech from the silent video of a talker using deep learning. The system learns a mapping function from raw video frames to acoustic features and reconstructs the speech with a vocoder synthesis algorithm. To improve speech reconstruction performance, our model is also trained to predict text information in a multi-task learning fashion and it is able to simultaneously reconstruct and recognise speech in real time. The results in terms of estimated speech quality and intelligibility show the effectiveness of our method, which exhibits an improvement over existing video-to-speech approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Task LearningSpeech SynthesisSimilar Papers 제목 키워드 기반
RobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting Self-Supervised Representations
Significant progress has been made in speaker dependent Lip-to-Speech synthesis, which aims to generate speech from silent videos of talking faces. Current state-of-the-art approaches primarily employ non-autoregressive …
Lip to Speech SynthesisSpeaker-Specific Lip to Speech SynthesisSpeech SynthesisIntelligible Lip-to-Speech Synthesis with Speech Units
In this paper, we propose a novel Lip-to-Speech synthesis (L2S) framework, for synthesizing intelligible speech from a silent lip movement video. Specifically, to complement the insufficient supervisory signal of the pre…
Lip to Speech SynthesisSpeech SynthesisFastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis
Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mainly use sequence-to-sequence models to s…
Lip to Speech SynthesisSpeech SynthesisSpeech Synthesis from Text and Ultrasound Tongue Image-based Articulatory Input
Articulatory information has been shown to be effective in improving the performance of HMM-based and DNN-based text-to-speech synthesis. Speech synthesis research focuses traditionally on text-to-speech conversion, when…
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisVisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection
The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly …
feature selectionSpeech Synthesis