Disjoint Mapping Network for Cross-modal Matching of Voices and Faces
We propose a novel framework, called Disjoint Mapping Network (DIMNet), for cross-modal biometric matching, in particular of voices and faces. Different from the existing methods, DIMNet does not explicitly learn the joint relationship between the modalities. Instead, DIMNet learns a shared representation for different modalities by mapping them individually to their common covariates. These shared representations can then be used to find the correspondences between the modalities. We show empirically that DIMNet is able to achieve better performance than other current methods, with the additional benefits of being conceptually simpler and less data-intensive.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
On Learning Associations of Faces and Voices
In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that t…
Speaker IdentificationSeeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision
The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from …
Lip ReadingSelf-Supervised LearningSpeaker RecognitionReconstructing faces from voices
Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from thei…
Face Reconstruction from Voice using Generative Adversarial Networks
Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from thei…
Face ReconstructionCross-Modal Perceptionist: Can Face Geometry be Gleaned from Voices?
This work digs into a root question in human perception: can face geometry be gleaned from one's voices? Previous works that study this question only adopt developments in image synthesis and convert voices into face ima…
3D Face Modelling3D Face ReconstructionImage Generation