paper-with-me

홈 › Papers

Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning

2025-01-26 · Qian Yang, Calbert Graham

Voice conversion (VC) modifies voice characteristics while preserving linguistic content. This paper presents the Stepback network, a novel model for converting speaker identity using non-parallel data. Unlike traditional VC methods that rely on parallel data, our approach leverages deep learning techniques to enhance disentanglement completion and linguistic content preservation. The Stepback network incorporates a dual flow of different domain data inputs and uses constraints with self-destructive amendments to optimize the content encoder. Extensive experiments show that our model significantly improves VC performance, reducing training costs while achieving high-quality voice conversion. The Stepback network's design offers a promising solution for advanced voice conversion tasks.

📄 PDF Abstract BibTeX arXiv:2501.15613

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementMulti-Task LearningVoice Conversion

Similar Papers 제목 키워드 기반

Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using $β$-VAE

2022-10-25 · Hui Lu, Disong Wang, Xixin Wu, Zhiyong Wu 외

We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to de…

DisentanglementRepresentation LearningSpeech Representation LearningVoice Conversion

StyleStream: Real-Time Zero-Shot Voice Style Conversion

2026-02-23 · Yisi Liu, Nicholas Lee, Gopala Anumanchipalli arxiv

Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic content from style. While prior work h…

NoiseVC: Towards High Quality Zero-Shot Voice Conversion

2021-04-13 · Shijun Wang, Damian Borth

Voice conversion (VC) is a task that transforms voice from target audio to source without losing linguistic contents, it is challenging especially when source and target speakers are unseen during training (zero-shot VC)…

DisentanglementQuantizationVocal Bursts Intensity PredictionVoice Conversion

Self-Supervised Representations for Singing Voice Conversion

2023-03-21 · Tejas Jayashankar, JiLong Wu, Leda Sari, David Kant 외

A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Ve…

DisentanglementVoice Conversion

Many-to-Many Voice Conversion based Feature Disentanglement using Variational Autoencoder

2021-07-11 · Manh Luong, Viet Anh Tran

Voice conversion is a challenging task which transforms the voice characteristics of a source speaker to a target speaker without changing linguistic content. Recently, there have been many works on many-to-many Voice Co…

DisentanglementVoice Conversion