paper-with-me

홈 › Papers

Learning Speaker-specific Lip-to-Speech Generation

2022-06-04 · Munender Varshney, Ravindra Yadav, Vinay P. Namboodiri, Rajesh M Hegde

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every speaker has a different accent and speaking style, which can be inferred from their visual and speech features. This work aims to understand the correlation/mapping between speech and the sequence of lip movement of individual speakers in an unconstrained and large vocabulary. We model the frame sequence as a prior to the transformer in an auto-encoder setting and learned a joint embedding that exploits temporal properties of both audio and video. We learn temporal synchronization using deep metric learning, which guides the decoder to generate speech in sync with input lip movements. The predictive posterior thus gives us the generated speech in speaker speaking style. We have trained our model on the Grid and Lip2Wav Chemistry lecture dataset to evaluate single speaker natural speech generation tasks from lip movement in an unconstrained natural setting. Extensive evaluation using various qualitative and quantitative metrics with human evaluation also shows that our method outperforms the Lip2Wav Chemistry dataset(large vocabulary in an unconstrained setting) by a good margin across almost all evaluation metrics and marginally outperforms the state-of-the-art on GRID dataset.

📄 PDF Abstract BibTeX arXiv:2206.02050

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderLip ReadingMetric Learning

Similar Papers 제목 키워드 기반

Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation

2025-01-11 · Zhengyan Sheng, Zhihao Du, Heng Lu, Shiliang Zhang 외

Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generation remains on the rise. This paper intr…

Diversity

We Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings

2024-07-05 · Ismail Rasim Ulgen, Carlos Busso, John H. L. Hansen, Berrak Sisman

Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, the…

Speaker RecognitionSpeech SynthesisVoice Conversion

CrossSpeech: Speaker-independent Acoustic Representation for Cross-lingual Speech Synthesis

2023-02-28 · Ji-Hoon Kim, Hong-Sun Yang, Yoon-Cheol Ju, Il-Hwan Kim 외

While recent text-to-speech (TTS) systems have made remarkable strides toward human-level quality, the performance of cross-lingual TTS lags behind that of intra-lingual TTS. This gap is mainly rooted from the speaker-la…

Speech Synthesistext-to-speechText to Speech

GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech Synthesis

2021-06-29 · Jinhyeok Yang, Jae-Sung Bae, Taejun Bak, Youngik Kim 외

Recent advances in neural multi-speaker text-to-speech (TTS) models have enabled the generation of reasonably good speech quality with a single model and made it possible to synthesize the speech of a speaker with limite…

Speech Synthesistext-to-speechText to SpeechVocal Bursts Intensity Prediction

CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation

2024-12-28 · Ji-Hoon Kim, Hong-Sun Yang, Yoon-Cheol Ju, Il-Hwan Kim 외

The goal of this work is to generate natural speech in multiple languages while maintaining the same speaker identity, a task known as cross-lingual speech synthesis. A key challenge of cross-lingual speech synthesis is …

Speech Synthesis