paper-with-me

홈 › Papers

DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

2024-06-13 · Neha Sahipjohn, Ashishkumar Gudmalwar, Nirmesh Shah, Pankaj Wasnik, Rajiv Ratn Shah

Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech (TTS), which can control the speech duration of synthesized speech in such a way that it aligns well with the speakers lip movements given in the reference video even when the spoken text is different or in a different language. To accomplish this, we propose to utilize cross-modal attention techniques in a pre-trained GPT-based TTS. We combine linguistic tokens from text, speaker identity tokens via a voice cloning network, and video tokens via a proposed duration controller network. We demonstrate the effectiveness of our system on the Lip2Wav-Chemistry and LRS2 datasets. Also, the proposed method achieves improved lip sync and naturalness compared to the SOTAs for the same language but different text (i.e., non-parallel) and the different language, different text (i.e., cross-lingual) scenarios.

📄 PDF Abstract BibTeX arXiv:2406.08802

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Modeltext-to-speechText to SpeechVoice Cloning

Similar Papers 제목 키워드 기반

VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing

2022-11-30 · Yihan Wu, Junliang Guo, Xu Tan, Chen Zhang 외

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation…

Machine TranslationSentencespeech-recognitionSpeech Recognition+2

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration

2024-12-11 · Haowei Lou, Helen Paik, Wen Hu, Lina Yao

Recent advancements in text-to-speech (TTS) systems, such as FastSpeech and StyleSpeech, have significantly improved speech generation quality. However, these models often rely on duration generated by external tools lik…

text-to-speechText to Speech

Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

2025-08-12 · Chaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong 외 arxiv

Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information …

Machine Translation

Expressive, Variable, and Controllable Duration Modelling in TTS

2022-06-28 · Ammar Abbas, Thomas Merritt, Alexis Moinet, Sri Karlapati 외

Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to relying on previous statistical parametric s…

Normalising FlowsSpeech Synthesistext-to-speechText to Speech

Isochrony-Controlled Speech-to-Text Translation: A study on translating from Sino-Tibetan to Indo-European Languages

2024-11-11 · Midia Yousefi, Yao Qian, Junkun Chen, Gang Wang 외

End-to-end speech translation (ST), which translates source language speech directly into target language text, has garnered significant attention in recent years. Many ST applications require strict length control to en…

DecoderMachine TranslationSpeech-to-TextSpeech-to-Text Translation+1