paper-with-me

홈 › Papers

Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features

2026-03-03 · Kyle Janse van Rensburg, Benjamin van Niekerk, Herman Kamper arxiv

How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have considered whether speech characteristics are captured within individual dimensions of SSL features. In this paper we specifically look at speaker information using PCA on utterance-averaged representations. For a range of SSL models, we find that the principal dimension that explains most variance encodes pitch and associated characteristics like gender. Other individual principal dimensions correlate with intensity, noise levels, the second formant, and higher frequency characteristics. We then use synthesis analyses to show that the dimensions for most characteristics are isolated from each other's influence. We further show that characteristics can be changed by manipulating the corresponding dimensions.

📄 PDF Abstract BibTeX arXiv:2603.03096

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Augmentation adversarial training for self-supervised speaker recognition

2020-07-23 · Jaesung Huh, Hee Soo Heo, Jingu Kang, Shinji Watanabe 외

The goal of this work is to train robust speaker recognition models without speaker labels. Recent works on unsupervised speaker representations are based on contrastive learning in which they encourage within-utterance …

Contrastive LearningSpeaker Recognition

Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model

2023-04-24 · Kenichi Fujita, Takanori Ashihara, Hiroki Kanagawa, Takafumi Moriya 외

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector…

RhythmSelf-Supervised LearningSpeech Synthesistext-to-speech+2

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

2020-10-21 · Jennifer Williams, Yi Zhao, Erica Cooper, Junichi Yamagishi

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not generalize well to unseen speakers or c…

speaker-diarizationSpeaker DiarizationSpeech Synthesis

ParrotTTS: Text-to-Speech synthesis by exploiting self-supervised representations

2023-03-01 · Neil Shah, Saiteja Kosgi, Vishal Tambrahalli, Neha Sahipjohn 외

We present ParrotTTS, a modularized text-to-speech synthesis model leveraging disentangled self-supervised speech representations. It can train a multi-speaker variant effectively using transcripts from a single speaker.…

Self-Supervised LearningSpeech Synthesistext-to-speechText to Speech+1

Speaker Normalization for Self-supervised Speech Emotion Recognition

2022-02-02 · Itai Gat, Hagai Aronowitz, Weizhong Zhu, Edmilson Morais 외

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristi…

Emotion RecognitionSpeech Emotion Recognition