paper-with-me

홈 › Papers

End-to-End Voice Conversion with Information Perturbation

2022-06-15 · Qicong Xie, Shan Yang, Yi Lei, Lei Xie, Dan Su

The ideal goal of voice conversion is to convert the source speaker's speech to sound naturally like the target speaker while maintaining the linguistic content and the prosody of the source speech. However, current approaches are insufficient to achieve comprehensive source prosody transfer and target speaker timbre preservation in the converted speech, and the quality of the converted speech is also unsatisfied due to the mismatch between the acoustic model and the vocoder. In this paper, we leverage the recent advances in information perturbation and propose a fully end-to-end approach to conduct high-quality voice conversion. We first adopt information perturbation to remove speaker-related information in the source speech to disentangle speaker timbre and linguistic content and thus the linguistic information is subsequently modeled by a content encoder. To better transfer the prosody of the source speech to the target, we particularly introduce a speaker-related pitch encoder which can maintain the general pitch pattern of the source speaker while flexibly modifying the pitch intensity of the generated speech. Finally, one-shot voice conversion is set up through continuous speaker space modeling. Experimental results indicate that the proposed end-to-end approach significantly outperforms the state-of-the-art models in terms of intelligibility, naturalness, and speaker similarity.

📄 PDF Abstract BibTeX arXiv:2206.07569

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features

2022-11-09 · Ziqian Ning, Qicong Xie, Pengcheng Zhu, Zhichao Wang 외

Voice conversion for highly expressive speech is challenging. Current approaches struggle with the balancing between speaker similarity, intelligibility and expressiveness. To address this problem, we propose Expressive-…

DecoderVoice Conversion

PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data

2025-06-01 · Songjun Cao, Qinghua Wu, Jie Chen, Jin Li 외

As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system comprises a content encoder and a speaker e…

Voice Conversion

SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers

2024-01-30 · Guangke Chen, Yedi Zhang, Fu Song, Ting Wang 외

Singing voice conversion (SVC) automates song covers by converting a source singing voice from a source singer into a new singing voice with the same lyrics and melody as the source, but sounds like being covered by the …

Voice Conversion

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching

2025-06-01 · Jialong Zuo, Shengpeng Ji, Minghui Fang, Mingze Li 외

Zero-Shot Voice Conversion (VC) aims to transform the source speaker's timbre into an arbitrary unseen one while retaining speech content. Most prior work focuses on preserving the source's prosody, while fine-grained ti…

RhythmStyle TransferVoice Conversion

SIG-VC: A Speaker Information Guided Zero-shot Voice Conversion System for Both Human Beings and Machines

2021-11-06 · Haozhe Zhang, Zexin Cai, Xiaoyi Qin, Ming Li

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people's attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method fo…

DisentanglementSpeaker VerificationVoice CloningVoice Conversion