SynCLR: A Synthesis Framework for Contrastive Learning of out-of-domain Speech Representations
Learning generalizable speech representations for unseen samples in different domains has been a challenge with ever increasing importance to date. Although contrastive learning has been a prominent class of representation learning approaches, the state-of-the-art (SOTA) contrastive learning methods were found to have limited ability for learning unseen out-of-domain speech representations. This paper presents SynCLR, a synthesis framework for contrastive learning of speech representations that can be generalized over unseen domains. Specifically, instead of using data augmentation approach, SynCLR employs data synthesis for multi-view generation. To ensure a highly-varied conditional speech distribution in view generation, we design a novel diffusion-based speech synthesizer. A new contrastive loss is also proposed to construct multiple embedding spaces, each of which preserves view-sensitive information to reduce domain reliance for a better disentanglement. Our experiments showed that SynCLR outperformed the SOTA contrastive learning methods with a 17.2\% relative reduction of EER in speaker verification tested on an unseen speech corpus, and considerably reduced 50.8\% relative FIDs in a challenging speech-to-image translation task given out-of-domain test speeches.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningData AugmentationDisentanglementRepresentation LearningSpeaker VerificationSpeech SynthesisTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Vision from Models Rivals Learning Vision from Data
We introduce SynCLR, a novel approach for learning visual representations exclusively from synthetic images and synthetic captions, without any real data. We synthesize a large dataset of image captions using LLMs, then …
Contrastive LearningImage Captioningimage-classificationImage Classification+2Self-supervised Context-aware Style Representation for Expressive Speech Synthesis
Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of lab…
Contrastive LearningDeep ClusteringExpressive Speech SynthesisRepresentation Learning+1Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastive learning-based TTS approach to transfe…
Contrastive LearningExpressive Speech SynthesisSpeech SynthesisCONCSS: Contrastive-based Context Comprehension for Dialogue-appropriate Prosody in Conversational Speech Synthesis
Conversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While previous methods have already delved into…
Contrastive LearningSelf-Supervised LearningSpeech SynthesisOmniDRCA: Parallel Speech-Text Foundation Model via Dual-Resolution Speech Representations and Contrastive Alignment
Recent studies on end-to-end speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing …
cross-modal alignmentQuestion AnsweringSpeech SynthesisText Generation