paper-with-me

홈 › Papers

Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping

2023-09-25 · Minki Kang, Wooseok Han, Eunho Yang

Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech. In this paper, we propose Face-StyleSpeech, a zero-shot Text-To-Speech (TTS) synthesis model that generates natural speech conditioned on a face image rather than reference speech. We hypothesize that learning entire prosodic features from a face image poses a significant challenge. To address this, our TTS model incorporates both face and prosody encoders. The prosody encoder is specifically designed to model speech style characteristics that are not fully captured by the face image, allowing the face encoder to focus on extracting speaker-specific features such as timbre. Experimental results demonstrate that Face-StyleSpeech effectively generates more natural speech from a face image than baselines, even for unseen faces. Samples are available on our demo page.

📄 PDF Abstract BibTeX arXiv:2311.05844

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models

2022-11-17 · Minki Kang, Dongchan Min, Sung Ju Hwang

There has been a significant progress in Text-To-Speech (TTS) synthesis technology in recent years, thanks to the advancement in neural generative modeling. However, existing methods on any-speaker adaptive TTS have achi…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation

2021-06-06 · Dongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju Hwang

With rapid progress in neural text-to-speech (TTS) models, personalized speech generation is now in high demand for many applications. For practical applicability, a TTS model should generate high-quality speech with onl…

text-to-speechText to Speech

StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech

2024-08-27 · Haowei Lou, Helen Paik, Wen Hu, Lina Yao

This paper introduces StyleSpeech, a novel Text-to-Speech~(TTS) system that enhances the naturalness and accuracy of synthesized speech. Building upon existing TTS technologies, StyleSpeech incorporates a unique Style De…

parameter-efficient fine-tuningtext-to-speechText to Speech

StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis

2023-12-19 · Xueyuan Chen, Xi Wang, Shaofei Zhang, Lei He 외

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-s…

DecoderSpeech Synthesis

Zero-shot personalized lip-to-speech synthesis with face image based voice control

2023-05-09 · Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. Howev…

Lip to Speech SynthesisRepresentation LearningSpeech Synthesis