paper-with-me

Papers

Reconstructing faces from voices

2019-05-25 · Yandong Wen, Rita Singh, Bhiksha Raj

Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from their voice. The task is designed to answer the question: given an audio clip spoken by an unseen person, can we picture a face that has as many common elements, or associations as possible with the speaker, in terms of identity? To address this problem, we propose a simple but effective computational framework based on generative adversarial networks (GANs). The network learns to generate faces from voices by matching the identities of generated faces to those of the speakers, on a training set. We evaluate the performance of the network by leveraging a closely related task - cross-modal matching. The results show that our model is able to generate faces that match several biometric characteristics of the speaker, and results in matching accuracies that are much better than chance.

📄 PDF Abstract BibTeX arXiv:1905.10604

Code (1)

cmu-mlsp/reconstructing_faces_from_voices pytorch

Similar Papers 제목 키워드 기반

Face Reconstruction from Voice using Generative Adversarial Networks

2019-12-01 · NeurIPS 2019 12 · Yandong Wen, Bhiksha Raj, Rita Singh

Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from thei…

Face Reconstruction

On Learning Associations of Faces and Voices

2018-05-15 · Changil Kim, Hijung Valentina Shin, Tae-Hyun Oh, Alexandre Kaspar 외

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that t…

Speaker Identification

PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association

2025-05-22 · Abdul Hannan, Muhammad Arslan Manzoor, Shah Nawaz, Muhammad Irzam Liaqat 외

We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as…

Voice2Mesh: Cross-Modal 3D Face Model Generation from Voices

2021-04-21 · Cho-Ying Wu, Ke Xu, Chin-Cheng Hsu, Ulrich Neumann

This work focuses on the analysis that whether 3D face models can be learned from only the speech inputs of speakers. Previous works for cross-modal face synthesis study image generation from voices. However, image synth…

Face GenerationFace ModelImage GenerationKnowledge Distillation

Naturalistic Music Decoding from EEG Data via Latent Diffusion Models

2024-05-15 · Emilian Postolache, Natalia Polouliakh, Hiroaki Kitano, Akima Connelly 외

In this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroencephalogram (EEG) recordings. Unlike sim…

channel selectionEEGElectroencephalogram (EEG)