paper-with-me

Papers

Cross-modal Speaker Verification and Recognition: A Multilingual Perspective

2020-04-28 · Muhammad Saad Saeed, Shah Nawaz, Pietro Morerio, Arif Mahmood, Ignazio Gallo, Muhammad Haroon Yousaf, Alessio Del Bue

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association between faces and voices across multiple languages spoken by the same set of persons. The aim of this paper is to answer two closely related questions: "Is face-voice association language independent?" and "Can a speaker be recognised irrespective of the spoken language?". These two questions are very important to understand effectiveness and to boost development of multilingual biometric systems. To answer them, we collected a Multilingual Audio-Visual dataset, containing human speech clips of $154$ identities with $3$ language annotations extracted from various videos uploaded online. Extensive experiments on the three splits of the proposed dataset have been performed to investigate and answer these novel research questions that clearly point out the relevance of the multilingual problem.

📄 PDF Abstract BibTeX arXiv:2004.13780

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker RecognitionSpeaker Verification

Similar Papers 제목 키워드 기반

On the influence of language similarity in non-target speaker verification trials

2025-06-03 · Paul M. Reuter, Michael Jessen

In this paper, we investigate the influence of language similarity in cross-lingual non-target speaker verification trials using a state-of-the-art speaker verification system, ECAPA-TDNN, trained on multilingual and mon…

Speaker Verification

L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification

2026-06-16 · Hyung-Seok Oh, Deok-Hyeon Cho, Seung-Bin Kim, Seong-Whan Lee arxiv

Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become entangled with linguistic characteristics, degrading generalization across languages…

Speaker Verification

Multilingual Audio-Visual Smartphone Dataset And Evaluation

2021-09-09 · Hareesh Mandalapu, Aravinda Reddy P N, Raghavendra Ramachandra, K Sreenivasa Rao 외

Smartphones have been employed with biometric-based verification systems to provide security in highly sensitive applications. Audio-visual biometrics are getting popular due to their usability, and also it will be chall…

Speaker Recognition

WeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition

2022-06-01 · LREC 2022 6 · Karen Jones, Kevin Walker, Christopher Caruso, Jonathan Wright 외

The WeCanTalk (WCT) Corpus is a new multi-language, multi-modal resource for speaker recognition. The corpus contains Cantonese, Mandarin and English telephony and video speech data from over 200 multilingual speakers lo…

Speaker Recognition

Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization

2024-07-25 · Ruijie Tao, Zhan Shi, Yidi Jiang, Duc-Tuan Truong 외

The human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as ``cross-modal speaker verification''. This task poses significant challenges du…

speaker-diarizationSpeaker DiarizationSpeaker Verification