paper-with-me

홈 › Papers

Voice-Face Cross-modal Matching and Retrieval: A Benchmark

2019-11-21 · Chuyuan Xiong, Deyuan Zhang, Tao Liu, Xiaoyong Du

Cross-modal associations between voice and face from a person can be learnt algorithmically, which can benefit a lot of applications. The problem can be defined as voice-face matching and retrieval tasks. Much research attention has been paid on these tasks recently. However, this research is still in the early stage. Test schemes based on random tuple mining tend to have low test confidence. Generalization ability of models can not be evaluated by small scale datasets. Performance metrics on various tasks are scarce. A benchmark for this problem needs to be established. In this paper, first, a framework based on comprehensive studies is proposed for voice-face matching and retrieval. It achieves state-of-the-art performance with various performance metrics on different tasks and with high test confidence on large scale datasets, which can be taken as a baseline for the follow-up research. In this framework, a voice anchored L2-Norm constrained metric space is proposed, and cross-modal embeddings are learned with CNN-based networks and triplet loss in the metric space. The embedding learning process can be more effective and efficient with this strategy. Different network structures of the framework and the cross language transfer abilities of the model are also analyzed. Second, a voice-face dataset (with 1.15M face data and 0.29M audio data) from Chinese speakers is constructed, and a convenient and quality controllable dataset collection tool is developed. The dataset and source code of the paper will be published together with this paper.

📄 PDF Abstract BibTeX arXiv:1911.09338

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalTriplet

Methods 이 논문이 사용한 방법론

Test 설명 없음
Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

A Benchmark for Voice-Face Cross-Modal Matching and Retrieval

2021-01-01 · Chuyuan Xiong, Deyuan Zhang, Tao Liu, Xiaoyong Du 외

Cross-modal associations between a person's voice and face can be learned algorithmically, and this is a useful functionality in many audio and visual applications. The problem can be defined as two tasks: voice-face mat…

Retrieval

Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder

2024-04-15 · Chong Peng, Liqiang He, Dan Su

Today, there have been many achievements in learning the association between voice and face. However, most previous work models rely on cosine similarity or L2 distance to evaluate the likeness of voices and faces follow…

Binary ClassificationContrastive LearningRetrieval

Face and Voice Cross-modal Association with Learning Convex Feature Embedding

2026-07-30 · Taewan Kim, Jiwoo Kang arxiv

Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Pre…

Cross-modal Face- and Voice-style Transfer

2023-02-27 · Naoya Takahashi, Mayank K. Singh, Yuki Mitsufuji

Image-to-image translation and voice conversion enable the generation of a new facial image and voice while maintaining some of the semantics such as a pose in an image and linguistic content in audio, respectively. They…

DiversityImage-to-Image TranslationOpen-Ended Question AnsweringStyle Transfer+2

Learnable PINs: Cross-Modal Embeddings for Person Identity

2018-05-02 · ECCV 2018 9 · Arsha Nagrani, Samuel Albanie, Andrew Zisserman

We propose and investigate an identity sensitive joint embedding of face and voice. Such an embedding enables cross-modal retrieval from voice to face and from face to voice. We make the following four contributions: fir…

Cross-Modal RetrievalRetrieval