You said that?
We present a method for generating a video of a talking face. The method takes as inputs: (i) still images of the target face, and (ii) an audio speech segment; and outputs a video of the target face lip synched with the audio. The method runs in real time and is applicable to faces and audio not seen at training time. To achieve this we propose an encoder-decoder CNN model that uses a joint embedding of the face and audio to generate synthesised talking face video frames. The model is trained on tens of hours of unlabelled videos. We also show results of re-dubbing videos using speech from a different person.
Code (1)
Tasks
DecoderUnconstrained Lip-synchronizationSimilar Papers 제목 키워드 기반
He Said, She Said: Gender in the ACL Anthology
Extractive email thread summarization: Can we do better than He Said She Said?
SAID: Accelerating Diffusion-Based Language Models via Scaffold-Aware Iterative Decoding
Diffusion large language models (DLLMs) enable non-autoregressive generation by iteratively denoising corrupted token sequences with bidirectional context. Despite their ability to update multiple positions in parallel, …
A Multimodal Simultaneous Interpretation Prototype: Who Said What
“Who said what” is essential for users to understand video streams that have more than one speaker, but conventional simultaneous interpretation systems merely present “what was said” in the form of subtitles. Because th…
SentenceTAGTranslationSAIDS: A Novel Approach for Sentiment Analysis Informed of Dialect and Sarcasm
Sentiment analysis becomes an essential part of every social network, as it enables decision-makers to know more about users' opinions in almost all life aspects. Despite its importance, there are multiple issues it enco…
Dialect IdentificationLanguage ModellingSarcasm DetectionSentence+3