paper-with-me

Papers

Speech2Face: Learning the Face Behind a Voice

2019-05-23 · CVPR 2019 6 · Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Wojciech Matusik

How much can we infer about a person's looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural network to perform this task using millions of natural Internet/YouTube videos of people speaking. During training, our model learns voice-face correlations that allow it to produce images that capture various physical attributes of the speakers such as age, gender and ethnicity. This is done in a self-supervised manner, by utilizing the natural co-occurrence of faces and speech in Internet videos, without the need to model attributes explicitly. We evaluate and numerically quantify how--and in what manner--our Speech2Face reconstructions, obtained directly from audio, resemble the true face images of the speakers.

📄 PDF Abstract BibTeX arXiv:1905.09773

Code (3)

Aryan05/Generative-Modelling-of-Images-from-Speech_Speech2Face tf
ravising-h/Speech2Face pytorch
saiteja-talluri/Speech2Face tf

Similar Papers 제목 키워드 기반

Crossmodal Voice Conversion

2019-04-09 · Hirokazu Kameoka, Kou Tanaka, Aaron Valero Puche, Yasunori Ohishi 외

Humans are able to imagine a person's voice from the person's appearance and imagine the person's appearance from his/her voice. In this paper, we make the first attempt to develop a method that can convert speech into a…

DecoderVoice Conversion

Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation

2025-07-25 · Fang Kang, Yin Cao, Haoyu Chen arxiv

Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more cha…

Talking Face GenerationFace Alignment

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis

2025-05-25 · Minsu Kim, Pingchuan Ma, Honglie Chen, Stavros Petridis 외

This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

2026-07-29 · Carlos Muñoz-Romero, Jose A. Gonzalez-Lopez arxiv

Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual information is available, e.g. for historical figures or video-game characters. …

Speech Synthesis

Zero-shot personalized lip-to-speech synthesis with face image based voice control

2023-05-09 · Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. Howev…

Lip to Speech SynthesisRepresentation LearningSpeech Synthesis