Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS transcends traditional TTS systems by interpreting facial cues and adjusting to emotional nuances without dependence on labeled datasets. To address sparse audio-visual-emotional data, the model is trained using LRS3, CREMA-D, and MELD datasets, demonstrating its adaptability. FEIM-TTS's unique capability to produce high-quality, speaker-agnostic speech makes it suitable for creating adaptable voices for virtual characters. Moreover, FEIM-TTS significantly enhances accessibility for individuals with visual impairments or those who have trouble seeing. By integrating emotional nuances into TTS, our model enables dynamic and engaging auditory experiences for webcomics, allowing visually impaired users to enjoy these narratives more fully. Comprehensive evaluation evidences its proficiency in modulating emotion and intensity, advancing emotional speech synthesis and accessibility. Samples are available at: https://feim-tts.github.io/.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotional Speech SynthesisSpeech Synthesistext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation
Existing facial expression recognition (FER) methods typically fine-tune a pre-trained visual encoder using discrete labels. However, this form of supervision limits to specify the emotional concept of different facial e…
Facial Expression RecognitionFacial Expression Recognition (FER)Disentangling Identity and Pose for Facial Expression Recognition
Facial expression recognition (FER) is a challenging problem because the expression component is always entangled with other irrelevant factors, such as identity and head pose. In this work, we propose an identity and po…
DecoderDisentanglementFace RecognitionFacial Expression Recognition+1Motion Transfer-Enhanced StyleGAN for Generating Diverse Macaque Facial Expressions
Generating animal faces using generative AI techniques is challenging because the available training images are limited both in quantity and variation, particularly for facial expressions across individuals. In this stud…
Data AugmentationImage EditingPhotorealistic Facial Expression Synthesis by the Conditional Difference Adversarial Autoencoder
Photorealistic facial expression synthesis from single face image can be widely applied to face recognition, data augmentation for emotion recognition or entertainment. This problem is challenging, in part due to a pauci…
Data AugmentationDecoderEmotion RecognitionFace RecognitionLearning Facial Representations from the Cycle-consistency of Face
Faces manifest large variations in many aspects, such as identity, expression, pose, and face styling. Therefore, it is a great challenge to disentangle and extract these characteristics from facial images, especially in…
Face ReconstructionFacial Expression RecognitionFacial Expression Recognition (FER)Image-to-Image Translation+1