paper-with-me

홈 › Papers

Towards the Next Frontier in Speech Representation Learning Using Disentanglement

2024-07-02 · Varun Krishna, Sriram Ganapathy

The popular frameworks for self-supervised learning of speech representations have largely focused on frame-level masked prediction of speech regions. While this has shown promising downstream task performance for speech recognition and related tasks, this has largely ignored factors of speech that are encoded at coarser level, like characteristics of the speaker or channel that remain consistent through-out a speech utterance. In this work, we propose a framework for Learning Disentangled Self Supervised (termed as Learn2Diss) representations of speech, which consists of frame-level and an utterance-level encoder modules. The two encoders are initially learned independently, where the frame-level model is largely inspired by existing self supervision techniques, thereby learning pseudo-phonemic representations, while the utterance-level encoder is inspired by constrastive learning of pooled embeddings, thereby learning pseudo-speaker representations. The joint learning of these two modules consists of disentangling the two encoders using a mutual information based criterion. With several downstream evaluation experiments, we show that the proposed Learn2Diss achieves state-of-the-art results on a variety of tasks, with the frame-level encoder representations improving semantic tasks, while the utterance-level representations improve non-semantic tasks.

📄 PDF Abstract BibTeX arXiv:2407.02543

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementRepresentation LearningSelf-Supervised Learningspeech-recognitionSpeech RecognitionSpeech Representation Learning

Similar Papers 제목 키워드 기반

Investigating Speaker Embedding Disentanglement on Natural Read Speech

2023-08-08 · Michael Kuhlmann, Adrian Meise, Fritz Seebauer, Petra Wagner 외

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainabi…

DisentanglementFairnessRepresentation Learning

Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using $β$-VAE

2022-10-25 · Hui Lu, Disong Wang, Xixin Wu, Zhiyong Wu 외

We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to de…

DisentanglementRepresentation LearningSpeech Representation LearningVoice Conversion

Are disentangled representations all you need to build speaker anonymization systems?

2022-08-22 · Pierre Champion, Denis Jouvet, Anthony Larcher

Speech signals contain a lot of sensitive information, such as the speaker's identity, which raises privacy concerns when speech data get collected. Speaker anonymization aims to transform a speech signal to remove the s…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Disentanglement+5

Disentangled Speaker Representation Learning via Mutual Information Minimization

2022-08-17 · Sung Hwan Mun, Min Hyun Han, Minchan Kim, Dongjune Lee 외

Domain mismatch problem caused by speaker-unrelated feature has been a major topic in speaker recognition. In this paper, we propose an explicit disentanglement framework to unravel speaker-relevant features from speaker…

DisentanglementRepresentation LearningSpeaker RecognitionSpeaker Verification+1

VQMIVC: Vector Quantization and Mutual Information-Based Unsupervised Speech Representation Disentanglement for One-shot Voice Conversion

2021-06-18 · Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen 외

One-shot voice conversion (VC), which performs conversion across arbitrary speakers with only a single target-speaker utterance for reference, can be effectively achieved by speech representation disentanglement. Existin…

DisentanglementQuantizationVoice Conversion