paper-with-me

Papers

Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?

2022-04-27 · Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Zhuo Chen, Peidong Wang, Gang Liu, Jinyu Li, Jian Wu, Xiangzhan Yu, Furu Wei

Recently, self-supervised learning (SSL) has demonstrated strong performance in speaker recognition, even if the pre-training objective is designed for speech recognition. In this paper, we study which factor leads to the success of self-supervised learning on speaker-related tasks, e.g. speaker verification (SV), through a series of carefully designed experiments. Our empirical results on the Voxceleb-1 dataset suggest that the benefit of SSL to SV task is from a combination of mask speech prediction loss, data scale, and model size, while the SSL quantizer has a minor impact. We further employ the integrated gradients attribution method and loss landscape visualization to understand the effectiveness of self-supervised learning for speaker recognition performance.

📄 PDF Abstract BibTeX arXiv:2204.12765

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeaker RecognitionSpeaker Verificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Self-supervised Learning with Random-projection Quantizer for Speech Recognition

2022-02-03 · Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu 외

We present a simple and effective self-supervised learning approach for speech recognition. The approach learns a model to predict the masked speech signals, in the form of discrete labels generated with a random-project…

Self-Supervised Learningspeech-recognitionSpeech Recognition

Towards End-to-end Unsupervised Speech Recognition

2022-04-05 · Alexander H. Liu, Wei-Ning Hsu, Michael Auli, Alexei Baevski

Unsupervised speech recognition has shown great potential to make Automatic Speech Recognition (ASR) systems accessible to every language. However, existing methods still heavily rely on hand-crafted pre-processing. Simi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?

2020-05-04 · Abhinav Shukla, Stavros Petridis, Maja Pantic

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between au…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionFace Reconstruction+5

Wav2Seq: Pre-training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages

2022-05-02 · Felix Wu, Kwangyoun Kim, Shinji Watanabe, Kyu Han 외

We introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervise…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decodernamed-entity-recognition+7

Recognizing More Emotions with Less Data Using Self-supervised Transfer Learning

2020-11-11 · Jonathan Boigne, Biman Liyanage, Ted Östrem

We propose a novel transfer learning method for speech emotion recognition allowing us to obtain promising results when only few training data is available. With as low as 125 examples per emotion class, we were able to …

Emotion RecognitionSpeech Emotion RecognitionTransfer Learning