paper-with-me

홈 › Papers

VoxBlink: A Large Scale Speaker Verification Dataset on Camera

2023-08-14 · Yuke Lin, Xiaoyi Qin, Guoqing Zhao, Ming Cheng, Ning Jiang, Haiyang Wu, Ming Li

In this paper, we introduce a large-scale and high-quality audio-visual speaker verification dataset, named VoxBlink. We propose an innovative and robust automatic audio-visual data mining pipeline to curate this dataset, which contains 1.45M utterances from 38K speakers. Due to the inherent nature of automated data collection, introducing noisy data is inevitable. Therefore, we also utilize a multi-modal purification step to generate a cleaner version of the VoxBlink, named VoxBlink-clean, comprising 18K identities and 1.02M utterances. In contrast to the VoxCeleb, the VoxBlink sources from short videos of ordinary users, and the covered scenarios can better align with real-life situations. To our best knowledge, the VoxBlink dataset is one of the largest publicly available speaker verification datasets. Leveraging the VoxCeleb and VoxBlink-clean datasets together, we employ diverse speaker verification models with multiple architectural backbones to conduct comprehensive evaluations on the VoxCeleb test sets. Experimental results indicate a substantial enhancement in performance,ranging from 12% to 30% relatively, across various backbone architectures upon incorporating the VoxBlink-clean into the training process. The details of the dataset can be found on http://voxblink.github.io

📄 PDF Abstract BibTeX arXiv:2308.07056

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker RecognitionSpeaker Verification

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

2024-07-16 · Yuke Lin, Ming Cheng, FuLin Zhang, Yingying Gao 외

In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This dataset represents a significant expans…

DiversitySpeaker IdentificationSpeaker RecognitionSpeaker Verification

The DKU-MSXF Speaker Verification System for the VoxCeleb Speaker Recognition Challenge 2023

2023-08-17 · Ze Li, Yuke Lin, Xiaoyi Qin, Ning Jiang 외

This paper is the system description of the DKU-MSXF System for the track1, track2 and track3 of the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). For Track 1, we utilize a network structure based on ResNet fo…

Domain AdaptationSemi-supervised Domain AdaptationSpeaker RecognitionSpeaker Verification

Analysis of ABC Frontend Audio Systems for the NIST-SRE24

2025-05-21 · Sara Barahona, Anna Silnova, Ladislav Mošner, Junyi Peng 외

We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of teleph…

Speaker Recognition

A Multi Purpose and Large Scale Speech Corpus in Persian and English for Speaker and Speech Recognition: the DeepMine Database

2019-12-08 · Hossein Zeinali, Lukáš Burget, Jan "Honza'' Černocký

DeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains mor…

Speaker Verificationspeech-recognitionSpeech RecognitionText-Dependent Speaker Verification+1

VoxAging: Continuously Tracking Speaker Aging with a Large-Scale Longitudinal Dataset in English and Mandarin

2025-05-27 · Zhiqi Ai, Meixuan Bao, Zhiyong Chen, Zhi Yang 외

The performance of speaker verification systems is adversely affected by speaker aging. However, due to challenges in data collection, particularly the lack of sustained and large-scale longitudinal data for individuals,…

Speaker Verification