paper-with-me

홈 › Papers

Audio-visual video-to-speech synthesis with synthesized input audio

2023-07-31 · Triantafyllos Kefalas, Yannis Panagakis, Maja Pantic

Video-to-speech synthesis involves reconstructing the speech signal of a speaker from a silent video. The implicit assumption of this task is that the sound signal is either missing or contains a high amount of noise/corruption such that it is not useful for processing. Previous works in the literature either use video inputs only or employ both video and audio inputs during training, and discard the input audio pathway during inference. In this work we investigate the effect of using video and audio inputs for video-to-speech synthesis during both training and inference. In particular, we use pre-trained video-to-speech models to synthesize the missing speech signals and then train an audio-visual-to-speech synthesis model, using both the silent video and the synthesized speech as inputs, to predict the final reconstructed speech. Our experiments demonstrate that this approach is successful with both raw waveforms and mel spectrograms as target outputs.

📄 PDF Abstract BibTeX arXiv:2307.16584

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Audiovisual Speech Synthesis using Tacotron2

2020-08-03 · Ahmed Hussen Abdelaziz, Anushree Prasanna Kumar, Chloe Seivwright, Gabriele Fanelli 외

Audiovisual speech synthesis is the problem of synthesizing a talking face while maximizing the coherency of the acoustic and visual speech. In this paper, we propose and compare two audiovisual speech synthesis systems …

Face ModelSentenceSpeech Synthesis

VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models

2025-04-03 · Kim Sung-Bin, Jeongsoo Choi, Puyuan Peng, Joon Son Chung 외

We present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting v…

Speech Synthesis

FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis

2025-03-06 · Ziqi Ni, Ao Fu, Yi Zhou

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computa…

Audio-Visual Synchronization

Large-scale multilingual audio visual dubbing

2020-11-06 · Yi Yang, Brendan Shillingford, Yannis Assael, Miaosen Wang 외

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed to text, translated, and automatically s…

Translation

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

2026-07-29 · Carlos Muñoz-Romero, Jose A. Gonzalez-Lopez arxiv

Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. …

Speech Synthesis