paper-with-me

홈 › Papers

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

2024-09-16 · Zakaria Aldeneh, Takuya Higuchi, Jee-weon Jung, Li-Wei Chen, Stephen Shum, Ahmed Hussen Abdelaziz, Shinji Watanabe, Tatiana Likhomanenko, Barry-John Theobald

Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker recognition start with representations extracted from very elaborate self-supervised methods (e.g., DINO). However, training such strong self-supervised models is not straightforward (they require hyper-parameter tuning and may not generalize to out-of-domain data) and, moreover, may not be needed at all. To this end, we show that the simple, well-studied, and established i-vector generative model is enough to bootstrap the IPL process for the unsupervised learning of speaker representations. We also systematically study the impact of other components on the IPL process, which includes the initial model, the encoder, augmentations, the number of clusters, and the clustering algorithm. Remarkably, we find that even with a simple and significantly weaker initial model like i-vector, IPL can still achieve speaker verification performance that rivals state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2409.10791

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker RecognitionSpeaker Verification

Methods 이 논문이 사용한 방법론

IPL Iterative Pseudo-Labeling (IPL) is a semi-supervised algorithm for speech recognition which efficiently performs multiple iterations of pseudo-labeling on unlabeled data as…

Similar Papers 제목 키워드 기반

A Study of F0 Modification for X-Vector Based Speech Pseudonymization Across Gender

2021-01-21 · Pierre Champion, Denis Jouvet, Anthony Larcher

Speech pseudonymization aims at altering a speech signal to map the identifiable personal characteristics of a given speaker to another identity. In other words, it aims to hide the source speaker identity while preservi…

Neural Predictive Coding using Convolutional Neural Networks towards Unsupervised Learning of Speaker Characteristics

2018-02-22 · Arindam Jati, Panayiotis Georgiou

Learning speaker-specific features is vital in many applications like speaker recognition, diarization and speech recognition. This paper provides a novel approach, we term Neural Predictive Coding (NPC), to learn speake…

Speaker IdentificationSpeaker RecognitionSpeaker Verificationspeech-recognition+1

Design Choices for X-vector Based Speaker Anonymization

2020-05-18 · Brij Mohan Lal Srivastava, Natalia Tomashenko, Xin Wang, Emmanuel Vincent 외

The recently proposed x-vector based anonymization scheme converts any input voice into that of a random pseudo-speaker. In this paper, we present a flexible pseudo-speaker selection technique as a baseline for the first…

Speaker anonymizationSpeaker Verification

AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios

2022-04-01 · Yihan Wu, Xu Tan, Bohan Li, Lei He 외

Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen spe…

Speech Synthesistext-to-speechText to Speech

The Phonexia VoxCeleb Speaker Recognition Challenge 2021 System Description

2021-09-05 · Josef Slavíček, Albert Swart, Michal Klčo, Niko Brümmer

We describe the Phonexia submission for the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21) in the unsupervised speaker verification track. Our solution was very similar to IDLab's winning submission for VoxSRC-2…

ClusteringContrastive LearningSpeaker RecognitionSpeaker Verification