paper-with-me

홈 › Papers

Robust Speaker Recognition with Transformers Using wav2vec 2.0

2022-03-28 · Sergey Novoselov, Galina Lavrentyeva, Anastasia Avdeeva, Vladimir Volokhov, Aleksei Gusev

Recent advances in unsupervised speech representation learning discover new approaches and provide new state-of-the-art for diverse types of speech processing tasks. This paper presents an investigation of using wav2vec 2.0 deep speech representations for the speaker recognition task. The proposed fine-tuning procedure of wav2vec 2.0 with simple TDNN and statistic pooling back-end using additive angular margin loss allows to obtain deep speaker embedding extractor that is well-generalized across different domains. It is concluded that Contrastive Predictive Coding pretraining scheme efficiently utilizes the power of unlabeled data, and thus opens the door to powerful transformer-based speaker recognition systems. The experimental results obtained in this study demonstrate that fine-tuning can be done on relatively small sets and a clean version of data. Using data augmentation during fine-tuning provides additional performance gains in speaker verification. In this study speaker recognition systems were analyzed on a wide range of well-known verification protocols: VoxCeleb1 cleaned test set, NIST SRE 18 development set, NIST SRE 2016 and NIST SRE 2019 evaluation set, VOiCES evaluation set, NIST 2021 SRE, and CTS challenges sets.

📄 PDF Abstract BibTeX arXiv:2203.15095

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationRepresentation LearningSpeaker RecognitionSpeaker VerificationSpeech Representation Learning

Methods 이 논문이 사용한 방법론

InfoNCE 설명 없음
Contrastive Predictive Coding Contrastive Predictive Coding (CPC) learns self-supervised representations by predicting the future in latent space by using powerful autoregressive models. The model uses a…

Similar Papers 제목 키워드 기반

SPEAKER VGG CCT: Cross-corpus Speech Emotion Recognition with Speaker Embedding and Vision Transformers

2022-11-04 · A. Arezzo, S. Berretti

In recent years, Speech Emotion Recognition (SER) has been investigated mainly transforming the speech signal into spectrograms that are then classified using Convolutional Neural Networks pretrained on generic images an…

Cross-corpusEmotion RecognitionSpeech Emotion Recognition

Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings

2025-03-13 · Jakaria Islam Emon, Md Abu Salek, Kazi Tamanna Alam

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification),…

Speaker Identificationspeech-recognitionSpeech Recognition

Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization

2022-08-27 · Dongmei Wang, Xiong Xiao, Naoyuki Kanda, Takuya Yoshioka 외

This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model's drawback of being unable to handle an arbitrary numbe…

Action DetectionActivity DetectionDecoderspeaker-diarization+1

AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers

2024-06-13 · Emil Biju, Anirudh Sriram, Mert Pilanci

While large transformer-based models have exhibited remarkable performance in speaker-independent speech recognition, their large size and computational requirements make them expensive or impractical to use in resource-…

speech-recognitionSpeech Recognition

PatchGame: Learning to Signal Mid-level Patches in Referential Games

2021-11-02 · NeurIPS 2021 12 · Kamal Gupta, Gowthami Somepalli, Anubhav Gupta, Vinoj Jayasundara 외

We study a referential game (a type of signaling game) where two agents communicate with each other via a discrete bottleneck to achieve a common goal. In our referential game, the goal of the speaker is to compose a mes…