paper-with-me

Papers

Speaker Recognition in the Wild

2022-05-05 · Neeraj Chhimwal, Anirudh Gupta, Rishabh Gaur, Harveen Singh Chadha, Priyanshi Shah, Ankur Dhuriya, Vivek Raghavan

In this paper, we propose a pipeline to find the number of speakers, as well as audios belonging to each of these now identified speakers in a source of audio data where number of speakers or speaker labels are not known a priori. We used this approach as a part of our Data Preparation pipeline for Speech Recognition in Indic Languages (https://github.com/Open-Speech-EkStep/vakyansh-wav2vec2-experimentation). To understand and evaluate the accuracy of our proposed pipeline, we introduce two metrics: Cluster Purity, and Cluster Uniqueness. Cluster Purity quantifies how "pure" a cluster is. Cluster Uniqueness, on the other hand, quantifies what percentage of clusters belong only to a single dominant speaker. We discuss more on these metrics in section \ref{sec:metrics}. Since we develop this utility to aid us in identifying data based on speaker IDs before training an Automatic Speech Recognition (ASR) model, and since most of this data takes considerable effort to scrape, we also conclude that 98\% of data gets mapped to the top 80\% of clusters (computed by removing any clusters with less than a fixed number of utterances -- we do this to get rid of some very small clusters and use this threshold as 30), in the test set chosen.

📄 PDF Abstract BibTeX arXiv:2205.02475

Code (1)

Open-Speech-EkStep/vakyansh-wav2vec2-experimentation 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Utterance-level Aggregation For Speaker Recognition In The Wild

2019-02-26 · Weidi Xie, Arsha Nagrani, Joon Son Chung, Andrew Zisserman

The objective of this paper is speaker recognition "in the wild"-where utterances may be of variable length and also contain irrelevant signals. Crucial elements in the design of deep networks for this task are the type …

Speaker RecognitionText-Independent Speaker Verification

VoxSRC 2019: The first VoxCeleb Speaker Recognition Challenge

2019-12-05 · Joon Son Chung, Arsha Nagrani, Ernesto Coto, Weidi Xie 외

The VoxCeleb Speaker Recognition Challenge 2019 aimed to assess how well current speaker recognition technology is able to identify speakers in unconstrained or `in the wild' data. It consisted of: (i) a publicly availab…

Speaker Recognition

VoxSRC 2020: The Second VoxCeleb Speaker Recognition Challenge

2020-12-12 · Arsha Nagrani, Joon Son Chung, Jaesung Huh, Andrew Brown 외

We held the second installment of the VoxCeleb Speaker Recognition Challenge in conjunction with Interspeech 2020. The goal of this challenge was to assess how well current speaker recognition technology is able to diari…

Speaker Recognition

Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition

2023-09-21 · Shuai Wang, Qibing Bai, Qi Liu, Jianwei Yu 외

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to tr…

Speaker Recognition

VoxSRC 2022: The Fourth VoxCeleb Speaker Recognition Challenge

2023-02-20 · Jaesung Huh, Andrew Brown, Jee-weon Jung, Joon Son Chung 외

This paper summarises the findings from the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22), which was held in conjunction with INTERSPEECH 2022. The goal of this challenge was to evaluate how well state-of-the-a…

Speaker DiarizationSpeaker RecognitionSpeaker Verification