Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning
Voice conversion (VC) modifies voice characteristics while preserving linguistic content. This paper presents the Stepback network, a novel model for converting speaker identity using non-parallel data. Unlike traditional VC methods that rely on parallel data, our approach leverages deep learning techniques to enhance disentanglement completion and linguistic content preservation. The Stepback network incorporates a dual flow of different domain data inputs and uses constraints with self-destructive amendments to optimize the content encoder. Extensive experiments show that our model significantly improves VC performance, reducing training costs while achieving high-quality voice conversion. The Stepback network's design offers a promising solution for advanced voice conversion tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
DisentanglementMulti-Task LearningVoice ConversionSimilar Papers 제목 키워드 기반
Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using $β$-VAE
We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to de…
DisentanglementRepresentation LearningSpeech Representation LearningVoice ConversionStyleStream: Real-Time Zero-Shot Voice Style Conversion
Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic content from style. While prior work h…
NoiseVC: Towards High Quality Zero-Shot Voice Conversion
Voice conversion (VC) is a task that transforms voice from target audio to source without losing linguistic contents, it is challenging especially when source and target speakers are unseen during training (zero-shot VC)…
DisentanglementQuantizationVocal Bursts Intensity PredictionVoice ConversionSelf-Supervised Representations for Singing Voice Conversion
A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Ve…
DisentanglementVoice ConversionMany-to-Many Voice Conversion based Feature Disentanglement using Variational Autoencoder
Voice conversion is a challenging task which transforms the voice characteristics of a source speaker to a target speaker without changing linguistic content. Recently, there have been many works on many-to-many Voice Co…
DisentanglementVoice Conversion