paper-with-me

Papers

Improved Speech Reconstruction from Silent Video

2017-08-01 · Ariel Ephrat, Tavi Halperin, Shmuel Peleg

Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible and natural-sounding acoustic speech signal from silent video frames of a speaking person. We train our model on speakers from the GRID and TCD-TIMIT datasets, and evaluate the quality and intelligibility of reconstructed speech using common objective measurements. We show that speech predictions from the proposed model attain scores which indicate significantly improved quality over existing models. In addition, we show promising results towards reconstructing speech from an unconstrained dictionary.

📄 PDF Abstract BibTeX arXiv:1708.01204

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vocoder-Based Speech Synthesis from Silent Videos

2020-04-06 · Daniel Michelsanti, Olga Slizovskaia, Gloria Haro, Emilia Gómez 외

Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this pape…

Multi-Task LearningSpeech Synthesis

Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training

2023-04-12 · Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

This paper studies the task of speech reconstruction from ultrasound tongue images and optical lip videos recorded in a silent speaking mode, where people only activate their intra-oral and extra-oral articulators withou…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Vid2speech: Speech Reconstruction from Silent Video

2017-01-02 · Ariel Ephrat, Shmuel Peleg

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from s…

Lipper: Synthesizing Thy Speech using Multi-View Lipreading

2019-06-28 · Yaman Kumar, Rohit Jain, Khwaja Mohd. Salik, Rajiv Ratn Shah 외

Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos int…

Lipreading

Lip2AudSpec: Speech reconstruction from silent lip movements video

2017-10-26 · Hassan Akbari, Himani Arora, Liangliang Cao, Nima Mesgarani

In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of speech and its corresponding sound generat…

Lip Reading