paper-with-me

Papers

Contrastive Speaker Embedding With Sequential Disentanglement

2023-09-23 · Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien

Contrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker identity but also linguistic content. In this paper, we propose a contrastive learning framework with sequential disentanglement to remove linguistic content by incorporating a disentangled sequential variational autoencoder (DSVAE) into the conventional SimCLR framework. The DSVAE aims to disentangle speaker factors from content factors in an embedding space so that only the speaker factors are used for constructing a contrastive loss objective. Because content factors have been removed from the contrastive learning, the resulting speaker embeddings will be content-invariant. Experimental results on VoxCeleb1-test show that the proposed method consistently outperforms SimCLR. This suggests that applying sequential disentanglement is beneficial to learning speaker-discriminative embeddings.

📄 PDF Abstract BibTeX arXiv:2309.13253

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDisentanglement

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Average Pooling 설명 없음
Batch Normalization 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Residual Connection 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…

Similar Papers 제목 키워드 기반

Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder

2024-09-05 · Yuying Xie, Michael Kuhlmann, Frederik Rautenberg, Zheng-Hua Tan 외

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversio…

DisentanglementVoice Conversion

Disentangled Speaker Representation Learning via Mutual Information Minimization

2022-08-17 · Sung Hwan Mun, Min Hyun Han, Minchan Kim, Dongjune Lee 외

Domain mismatch problem caused by speaker-unrelated feature has been a major topic in speaker recognition. In this paper, we propose an explicit disentanglement framework to unravel speaker-relevant features from speaker…

DisentanglementRepresentation LearningSpeaker RecognitionSpeaker Verification+1

Investigating Speaker Embedding Disentanglement on Natural Read Speech

2023-08-08 · Michael Kuhlmann, Adrian Meise, Fritz Seebauer, Petra Wagner 외

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainabi…

DisentanglementFairnessRepresentation Learning

Robust Disentangled Variational Speech Representation Learning for Zero-shot Voice Conversion

2022-03-30 · Jiachen Lian, Chunlei Zhang, Dong Yu

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functi…

Data AugmentationDecoderDisentanglementRepresentation Learning+3

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis

2024-12-22 · Ye-Xin Lu, Hui-Peng Du, Zheng-Yan Sheng, Yang Ai 외

This paper proposes an Incremental Disentanglement-based Environment-Aware zero-shot text-to-speech (TTS) method, dubbed IDEA-TTS, that can synthesize speech for unseen speakers while preserving the acoustic characterist…

DecoderDisentanglementSpeech Synthesistext-to-speech+2