paper-with-me

홈 › Papers

SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

2023-12-14 · Junjie Li, Yiwei Guo, Xie Chen, Kai Yu

Zero-shot voice conversion (VC) aims to transfer the source speaker timbre to arbitrary unseen target speaker timbre, while keeping the linguistic content unchanged. Although the voice of generated speech can be controlled by providing the speaker embedding of the target speaker, the speaker similarity still lags behind the ground truth recordings. In this paper, we propose SEF-VC, a speaker embedding free voice conversion model, which is designed to learn and incorporate speaker timbre from reference speech via a powerful position-agnostic cross-attention mechanism, and then reconstruct waveform from HuBERT semantic tokens in a non-autoregressive manner. The concise design of SEF-VC enhances its training stability and voice conversion performance. Objective and subjective evaluations demonstrate the superiority of SEF-VC to generate high-quality speech with better similarity to target reference than strong zero-shot VC baselines, even for very short reference speeches.

📄 PDF Abstract BibTeX arXiv:2312.08676

Code (0)

등록된 구현이 없습니다.

Tasks

PositionVoice Conversion

Similar Papers 제목 키워드 기반

Improving Zero-shot Voice Style Transfer via Disentangled Representation Learning

2021-03-17 · ICLR 2021 1 · Siyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao 외

Voice style transfer, also called voice conversion, seeks to modify one speaker's voice to generate speech as if it came from another (target) speaker. Previous works have made progress on voice conversion with parallel …

DecoderRepresentation LearningStyle TransferVoice Conversion

SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement

2024-07-10 · ZiHao Wang, Le Ma, Yongsheng Feng, Xin Pan 외

Singing voice conversion (SVC) aims to convert a singer's voice to another singer's from a reference audio while keeping the original semantics. However, existing SVC methods can hardly perform zero-shot due to incomplet…

DisentanglementVoice Conversion

Zero-shot personalized lip-to-speech synthesis with face image based voice control

2023-05-09 · Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. Howev…

Lip to Speech SynthesisRepresentation LearningSpeech Synthesis

Robust Disentangled Variational Speech Representation Learning for Zero-shot Voice Conversion

2022-03-30 · Jiachen Lian, Chunlei Zhang, Dong Yu

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functi…

Data AugmentationDecoderDisentanglementRepresentation Learning+3

StarGAN-ZSVC: Towards Zero-Shot Voice Conversion in Low-Resource Contexts

2021-05-31 · Matthew Baas, Herman Kamper

Voice conversion is the task of converting a spoken utterance from a source speaker so that it appears to be said by a different target speaker while retaining the linguistic content of the utterance. Recent advances hav…

Voice Conversion