paper-with-me

홈 › Papers

A Benchmark for Voice-Face Cross-Modal Matching and Retrieval

2021-01-01 · Chuyuan Xiong, Deyuan Zhang, Tao Liu, Xiaoyong Du, Jiankun Tian, Songyan Xue

Cross-modal associations between a person's voice and face can be learned algorithmically, and this is a useful functionality in many audio and visual applications. The problem can be defined as two tasks: voice-face matching and retrieval. Recently, this topic has attracted much research attention, but it is still in its early stages of development, and evaluation protocols and test schemes need to be more standardized. Performance metrics for different subtasks are also scarce, and a benchmark for this problem needs to be established. In this paper, a baseline evaluation framework is proposed for voice-face matching and retrieval tasks. Test confidence is analyzed, and a confidence interval for estimated accuracy is proposed. Various state-of-the-art performances with high test confidence are achieved on a series of subtasks using the baseline method (called TriNet) included in this framework. The source code will be published along with the paper. The results of this study can provide a basis for future research on voice-face cross-modal learning.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Similar Papers 제목 키워드 기반

Voice-Face Cross-modal Matching and Retrieval: A Benchmark

2019-11-21 · Chuyuan Xiong, Deyuan Zhang, Tao Liu, Xiaoyong Du

Cross-modal associations between voice and face from a person can be learnt algorithmically, which can benefit a lot of applications. The problem can be defined as voice-face matching and retrieval tasks. Much research a…

RetrievalTriplet

Cross-modal Face- and Voice-style Transfer

2023-02-27 · Naoya Takahashi, Mayank K. Singh, Yuki Mitsufuji

Image-to-image translation and voice conversion enable the generation of a new facial image and voice while maintaining some of the semantics such as a pose in an image and linguistic content in audio, respectively. They…

DiversityImage-to-Image TranslationOpen-Ended Question AnsweringStyle Transfer+2

Seeing Voices and Hearing Faces: Cross-modal biometric matching

2018-04-01 · CVPR 2018 6 · Arsha Nagrani, Samuel Albanie, Andrew Zisserman

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at ans…

Face RecognitionSpeaker Identification

Disjoint Mapping Network for Cross-modal Matching of Voices and Faces

2018-07-12 · ICLR 2019 5 · Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj 외

We propose a novel framework, called Disjoint Mapping Network (DIMNet), for cross-modal biometric matching, in particular of voices and faces. Different from the existing methods, DIMNet does not explicitly learn the joi…

Reconstructing faces from voices

2019-05-25 · Yandong Wen, Rita Singh, Bhiksha Raj

Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from thei…