paper-with-me

홈 › Papers

PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data

2025-06-01 · Songjun Cao, Qinghua Wu, Jie Chen, Jin Li, Long Ma

As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system comprises a content encoder and a speaker encoder. However, two types of mismatches arise: one for the inputs to the content encoder during training and inference, and another for the inputs to the speaker encoder. To address these mismatches, we propose a novel VC training method called \textit{PseudoVC} in this paper. First, we introduce an innovative information perturbation approach named \textit{Pseudo Conversion} to tackle the first mismatch problem. This approach leverages pretrained VC models to convert the source utterance into a perturbed utterance, which is fed into the content encoder during training. Second, we propose an approach termed \textit{Speaker Sampling} to resolve the second mismatch problem, which will substitute the input to the speaker encoder by another utterance from the same speaker during training. Experimental results demonstrate that our proposed \textit{Pseudo Conversion} outperforms previous information perturbation methods, and the overall \textit{PseudoVC} method surpasses publicly available VC models. Audio examples are available.

📄 PDF Abstract BibTeX arXiv:2506.01039

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

DeID-VC: Speaker De-identification via Zero-shot Pseudo Voice Conversion

2022-09-09 · Ruibin Yuan, Yuxuan Wu, Jacob Li, Jaxter Kim

The widespread adoption of speech-based online services raises security and privacy concerns regarding the data that they use and share. If the data were compromised, attackers could exploit user speech to bypass speaker…

De-identificationSpeaker VerificationVoice Conversion

Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion

2024-06-04 · RuiQi Li, Rongjie Huang, Yongqi Wang, Zhiqing Hong 외

Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal…

In-Context LearningLanguage ModelingLanguage ModellingRhythm+3

StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion

2024-08-05 · Zhichao Wang, Yuanzhe Chen, Xinsheng Wang, Lei Xie 외

StreamVoice has recently pushed the boundaries of zero-shot voice conversion (VC) in the streaming domain. It uses a streamable language model (LM) with a context-aware approach to convert semantic features from automati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modellingspeech-recognition+2

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

2026-04-14 · Qixi Zheng, Yuxiang Zhao, Tianrui Wang, Wenxi Chen 외 arxiv

Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building ze…

Voice Conversion

Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice Alignment

2023-09-18 · Zheng-Yan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua Ling

This paper presents a novel task, zero-shot voice conversion based on face images (zero-shot FaceVC), which aims at converting the voice characteristics of an utterance from any source speaker to a newly coming target sp…

Voice Conversion