Probing Speaker-specific Features in Speaker Representations
This study explores speaker-specific features encoded in speaker embeddings and intermediate layers of speech self-supervised learning (SSL) models. By utilising a probing method, we analyse features such as pitch, tempo, and energy across prominent speaker embedding models and speech SSL models, including HuBERT, WavLM, and Wav2vec 2.0. The results reveal that speaker embeddings like CAM++ excel in energy classification, while speech SSL models demonstrate superior performance across multiple features due to their hierarchical feature encoding. Intermediate layers effectively capture a mix of acoustic and para-linguistic information, with deeper layers refining these representations. This investigation provides insights into model design and highlights the potential of these representations for downstream applications, such as speaker verification and text-to-speech synthesis, while laying the groundwork for exploring additional features and advanced probing methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Self-Supervised LearningSpeaker VerificationSpeech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisSimilar Papers 제목 키워드 기반
Analyzing Speaker Information in Self-Supervised Models to Improve Zero-Resource Speech Processing
Contrastive predictive coding (CPC) aims to learn representations of speech by distinguishing future observations from a set of negative examples. Previous work has shown that linear classifiers trained on CPC features c…
Acoustic Unit DiscoveryLanguage ModelingLanguage ModellingSpeaker VerificationBeyond Decodability: Reconstructing Language Model Representations with an Encoding Probe
Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two limitations that we aim to solve with our new encoding probe approach…
Probing Deep Speaker Embeddings for Speaker-related Tasks
Deep speaker embeddings have shown promising results in speaker recognition, as well as in other speaker-related tasks. However, some issues are still under explored, for instance, the information encoded in these repres…
Speaker RecognitionSpeaker Verificationtext-to-speechText to SpeechAudio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation
For self-supervised speech processing, it is crucial to use pretrained models as speech representation extractors. In recent works, increasing the size of the model has been utilized in acoustic model training in order t…
Self-Supervised LearningSpeaker IdentificationSelf-supervised Predictive Coding Models Encode Speaker and Phonetic Information in Orthogonal Subspaces
Self-supervised speech representations are known to encode both speaker and phonetic information, but how they are distributed in the high-dimensional space remains largely unexplored. We hypothesize that they are encode…
Disentanglement