paper-with-me

홈 › Papers

Discrete Unit based Masking for Improving Disentanglement in Voice Conversion

2024-09-17 · Philip H. Lee, Ismail Rasim Ulgen, Berrak Sisman

Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker's identity from linguistic information is crucial. However, the disentanglement approaches used in these methods are limited as the speaker features depend on the phonetic content of the utterance, compromising disentanglement. This dependency is amplified with attention-based methods. To address this, we introduce a novel masking mechanism in the input before speaker encoding, masking certain discrete speech units that correspond highly with phoneme classes. Our work aims to reduce the phonetic dependency of speaker features by restricting access to some phonetic information. Furthermore, since our approach is at the input level, it is applicable to any encoder-decoder based VC framework. Our approach improves disentanglement and conversion performance across multiple VC methods, showing significant effectiveness, particularly in attention-based method, with 44% relative improvement in objective intelligibility.

📄 PDF Abstract BibTeX arXiv:2409.11560

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDisentanglementVoice Conversion

Similar Papers 제목 키워드 기반

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

2025-06-04 · Seymanur Aktı, Tuan Nam Nguyen, Alexander Waibel

Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework wi…

DisentanglementStyle TransferVoice Conversion

A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion

2021-11-03 · Benjamin van Niekerk, Marc-André Carbonneau, Julian Zaïdi, Mathew Baas 외

The goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we com…

Representation LearningVoice Conversion

Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning

2025-01-26 · Qian Yang, Calbert Graham

Voice conversion (VC) modifies voice characteristics while preserving linguistic content. This paper presents the Stepback network, a novel model for converting speaker identity using non-parallel data. Unlike traditiona…

DisentanglementMulti-Task LearningVoice Conversion

Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using $β$-VAE

2022-10-25 · Hui Lu, Disong Wang, Xixin Wu, Zhiyong Wu 외

We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to de…

DisentanglementRepresentation LearningSpeech Representation LearningVoice Conversion

StyleStream: Real-Time Zero-Shot Voice Style Conversion

2026-02-23 · Yisi Liu, Nicholas Lee, Gopala Anumanchipalli arxiv

Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic content from style. While prior work h…