paper-with-me

홈 › Papers

Let There Be Sound: Reconstructing High Quality Speech from Silent Videos

2023-08-29 · Ji-Hoon Kim, Jaehun Kim, Joon Son Chung

The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2) multiple speech variations, resulting in a mispronounced and over-smoothed speech. In this paper, we propose a novel lip-to-speech system that significantly improves the generation quality by alleviating the one-to-many mapping problem from multiple perspectives. Specifically, we incorporate (1) self-supervised speech representations to disambiguate homophenes, and (2) acoustic variance information to model diverse speech styles. Additionally, to better solve the aforementioned problem, we employ a flow based post-net which captures and refines the details of the generated speech. We perform extensive experiments on two datasets, and demonstrate that our method achieves the generation quality close to that of real human utterance, outperforming existing methods in terms of speech naturalness and intelligibility by a large margin. Synthesised samples are available at our demo page: https://mm.kaist.ac.kr/projects/LTBS.

📄 PDF Abstract BibTeX arXiv:2308.15256

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lip2AudSpec: Speech reconstruction from silent lip movements video

2017-10-26 · Hassan Akbari, Himani Arora, Liangliang Cao, Nima Mesgarani

In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of speech and its corresponding sound generat…

Lip Reading

Improved Speech Reconstruction from Silent Video

2017-08-01 · Ariel Ephrat, Tavi Halperin, Shmuel Peleg

Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model ba…

Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

2021-05-13 · Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 외

Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these t…

DecoderSpeech Synthesistext-to-speechText to Speech+1

Text-to-speech for the hearing impaired

2020-12-03 · Josef Schlittenlacher, Thomas Baer

Text-to-speech (TTS) systems offer the opportunity to compensate for a hearing loss at the source rather than correcting for it at the receiving end. This removes limitations such as time constraints for algorithms that …

text-to-speechText to SpeechTransfer Learning

When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds

2025-05-30 · Minsu Kang, Seolhee Lee, Choonghyeon Lee, Namhyun Cho

Human to non-human voice conversion (H2NH-VC) transforms human speech into animal or designed vocalizations. Unlike prior studies focused on dog-sounds and 16 or 22.05kHz audio transformation, this work addresses a broad…

Voice Conversion