paper-with-me

홈 › Papers

Phone and speaker spatial organization in self-supervised speech representations

2023-02-24 · Pablo Riera, Manuela Cerdeiro, Leonardo Pepino, Luciana Ferrer

Self-supervised representations of speech are currently being widely used for a large number of applications. Recently, some efforts have been made in trying to analyze the type of information present in each of these representations. Most such work uses downstream models to test whether the representations can be successfully used for a specific task. The downstream models, though, typically perform nonlinear operations on the representation extracting information that may not have been readily available in the original representation. In this work, we analyze the spatial organization of phone and speaker information in several state-of-the-art speech representations using methods that do not require a downstream model. We measure how different layers encode basic acoustic parameters such as formants and pitch using representation similarity analysis. Further, we study the extent to which each representation clusters the speech samples by phone or speaker classes using non-parametric statistical testing. Our results indicate that models represent these speech attributes differently depending on the target task used during pretraining.

📄 PDF Abstract BibTeX arXiv:2302.14055

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

2020-10-21 · Jennifer Williams, Yi Zhao, Erica Cooper, Junichi Yamagishi

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not generalize well to unseen speakers or c…

speaker-diarizationSpeaker DiarizationSpeech Synthesis

Self-supervised Predictive Coding Models Encode Speaker and Phonetic Information in Orthogonal Subspaces

2023-05-21 · Oli Liu, Hao Tang, Sharon Goldwater

Self-supervised speech representations are known to encode both speaker and phonetic information, but how they are distributed in the high-dimensional space remains largely unexplored. We hypothesize that they are encode…

Disentanglement

Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations

2024-06-13 · Mukhtar Mohamed, Oli Danyi Liu, Hao Tang, Sharon Goldwater

Self-supervised speech representations can hugely benefit downstream speech technologies, yet the properties that make them useful are still poorly understood. Two candidate properties related to the geometry of the repr…

Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models

2025-06-12 · Michele Gubian, Ioana Krehan, Oli Liu, James Kirby 외

Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. Here, we examine how wav2vec2 models train…

Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation

2020-05-18 · Po-Han Chi, Pei-Hung Chung, Tsung-Han Wu, Chun-Cheng Hsieh 외

For self-supervised speech processing, it is crucial to use pretrained models as speech representation extractors. In recent works, increasing the size of the model has been utilized in acoustic model training in order t…

Self-Supervised LearningSpeaker Identification