paper-with-me

Papers

Self-supervised Predictive Coding Models Encode Speaker and Phonetic Information in Orthogonal Subspaces

2023-05-21 · Oli Liu, Hao Tang, Sharon Goldwater

Self-supervised speech representations are known to encode both speaker and phonetic information, but how they are distributed in the high-dimensional space remains largely unexplored. We hypothesize that they are encoded in orthogonal subspaces, a property that lends itself to simple disentanglement. Applying principal component analysis to representations of two predictive coding models, we identify two subspaces that capture speaker and phonetic variances, and confirm that they are nearly orthogonal. Based on this property, we propose a new speaker normalization method which collapses the subspace that encodes speaker information, without requiring transcriptions. Probing experiments show that our method effectively eliminates speaker information and outperforms a previous baseline in phone discrimination tasks. Moreover, the approach generalizes and can be used to remove information of unseen speakers.

📄 PDF Abstract BibTeX arXiv:2305.12464

Code (0)

등록된 구현이 없습니다.

Tasks

Disentanglement

Similar Papers 제목 키워드 기반

Vector-Quantized Autoregressive Predictive Coding

2020-05-17 · Yu-An Chung, Hao Tang, James Glass

Autoregressive Predictive Coding (APC), as a self-supervised objective, has enjoyed success in learning representations from large amounts of unlabeled data, and the learned representations are rich for many downstream t…

Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder

2024-09-05 · Yuying Xie, Michael Kuhlmann, Frederik Rautenberg, Zheng-Hua Tan 외

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversio…

DisentanglementVoice Conversion

Neural Predictive Coding using Convolutional Neural Networks towards Unsupervised Learning of Speaker Characteristics

2018-02-22 · Arindam Jati, Panayiotis Georgiou

Learning speaker-specific features is vital in many applications like speaker recognition, diarization and speech recognition. This paper provides a novel approach, we term Neural Predictive Coding (NPC), to learn speake…

Speaker IdentificationSpeaker RecognitionSpeaker Verificationspeech-recognition+1

Speaker Group Encoding in Self-supervised Speech Recognition Models

2026-06-09 · Felix Herron, Solange Rossato Alexandre Allauzen, Benoit Favre, François Portet arxiv

We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, finetuned on speaker identification (SID), finetuned on automatic spee…

Speaker IdentificationSpeech Recognition

Learning Invariant World State Representations with Predictive Coding

2022-07-06 · Avi Ziskind, Sujeong Kim, Giedrius T. Burachas

Self-supervised learning methods overcome the key bottleneck for building more capable AI: limited availability of labeled data. However, one of the drawbacks of self-supervised architectures is that the representations …

DecoderSelf-Supervised Learning