From Speech-to-Speech Translation to Automatic Dubbing
We present enhancements to a speech-to-speech translation pipeline in order to perform automatic dubbing. Our architecture features neural machine translation generating output of preferred length, prosodic alignment of the translation with the original speech segments, neural text-to-speech with fine tuning of the duration of each utterance, and, finally, audio rendering to enriches text-to-speech output with background noise and reverberation extracted from the original audio. We report on a subjective evaluation of automatic dubbing of excerpts of TED Talks from English into Italian, which measures the perceived naturalness of automatic dubbing and the relative importance of each proposed enhancement.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSpeech-to-Speech Translationtext-to-speechText to SpeechTranslationSimilar Papers 제목 키워드 기반
Jointly Optimizing Translations and Speech Timing to Improve Isochrony in Automatic Dubbing
Automatic dubbing (AD) is the task of translating the original speech in a video into target language speech. The new target language speech should satisfy isochrony; that is, the new speech should be time aligned with t…
TranslationIsochrony-Aware Neural Machine Translation for Automatic Dubbing
We introduce the task of isochrony-aware machine translation which aims at generating translations suitable for dubbing. Dubbing of a spoken sentence requires transferring the content as well as the speech-pause structur…
Machine TranslationSentenceTranslationMachine Translation Verbosity Control for Automatic Dubbing
Automatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey …
Machine TranslationTranslationProsodic Alignment for off-screen automatic dubbing
The goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence. This entails isochrony, i.e., translating the original speech by also matching its prosodic structure into p…
Speech-to-Speech TranslationTranslationLingua: Addressing Scenarios for Live Interpretation and Automatic Dubbing
Lingua is an application developed for the Church of Jesus Christ of Latter-day Saints that performs both real-time interpretation of live speeches and automatic video dubbing (AVD). Like other AVD systems, it can perfor…
Translation