paper-with-me

홈 › Papers

Speaker recognition with two-step multi-modal deep cleansing

2022-10-28 · Ruijie Tao, Kong Aik Lee, Zhan Shi, Haizhou Li

Neural network-based speaker recognition has achieved significant improvement in recent years. A robust speaker representation learns meaningful knowledge from both hard and easy samples in the training set to achieve good performance. However, noisy samples (i.e., with wrong labels) in the training set induce confusion and cause the network to learn the incorrect representation. In this paper, we propose a two-step audio-visual deep cleansing framework to eliminate the effect of noisy labels in speaker representation learning. This framework contains a coarse-grained cleansing step to search for the peculiar samples, followed by a fine-grained cleansing step to filter out the noisy labels. Our study starts from an efficient audio-visual speaker recognition system, which achieves a close to perfect equal-error-rate (EER) of 0.01\%, 0.07\% and 0.13\% on the Vox-O, E and H test sets. With the proposed multi-modal cleansing mechanism, four different speaker recognition networks achieve an average improvement of 5.9\%. Code has been made available at: \textcolor{magenta}{\url{https://github.com/TaoRuijie/AVCleanse}}.

📄 PDF Abstract BibTeX arXiv:2210.15903

Code (1)

taoruijie/avcleanse 공식 구현 pytorch

Tasks

Representation LearningSpeaker RecognitionVocal Bursts Valence Prediction

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Global-Local GCN: Large-Scale Label Noise Cleansing for Face Recognition

2020-06-01 · CVPR 2020 6 · Yaobin Zhang, Weihong Deng, Mei Wang, Jiani Hu 외

In the field of face recognition, large-scale web-collected datasets are essential for learning discriminative representations, but they suffer from noisy identity labels, such as outliers and label flips. It is benefici…

Face RecognitionRepresentation Learning

DeepMSRF: A novel Deep Multimodal Speaker Recognition framework with Feature selection

2020-07-14 · Ehsan Asali, Farzan Shenavarmasouleh, Farid Ghareh Mohammadi, Prasanth Sengadu Suresh 외

For recognizing speakers in video streams, significant research studies have been made to obtain a rich machine learning model by extracting high-level speaker's features such as facial expression, emotion, and gender. H…

feature selectionSpeaker Recognition

Look, Listen and Learn - A Multimodal LSTM for Speaker Identification

2016-02-13 · Jimmy Ren, Yongtao Hu, Yu-Wing Tai, Chuan Wang 외

Speaker identification refers to the task of localizing the face of a person who has the same identity as the ongoing voice in a video. This task not only requires collective perception over both visual and auditory sign…

Speaker Identification

WeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition

2022-06-01 · LREC 2022 6 · Karen Jones, Kevin Walker, Christopher Caruso, Jonathan Wright 외

The WeCanTalk (WCT) Corpus is a new multi-language, multi-modal resource for speaker recognition. The corpus contains Cantonese, Mandarin and English telephony and video speech data from over 200 multilingual speakers lo…

Speaker Recognition

Speaker Recognition in Realistic Scenario Using Multimodal Data

2023-02-25 · Saqlain Hussain Shah, Muhammad Saad Saeed, Shah Nawaz, Muhammad Haroon Yousaf

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental i…

Speaker Recognition