paper-with-me

Papers

PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

2023-09-17 · Jixun Yao, Yuguang Yang, Yi Lei, Ziqian Ning, Yanni Hu, Yu Pan, JingJing Yin, Hongbin Zhou, Heng Lu, Lei Xie

Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to control the conversion process, which leads to limitations in style diversity or falls short in terms of the intuitive and interpretability of style representation. In this study, we propose PromptVC, a novel style voice conversion approach that employs a latent diffusion model to generate a style vector driven by natural language prompts. Specifically, the style vector is extracted by a style encoder during training, and then the latent diffusion model is trained independently to sample the style vector from noise, with this process being conditioned on natural language prompts. To improve style expressiveness, we leverage HuBERT to extract discrete tokens and replace them with the K-Means center embedding to serve as the linguistic content, which minimizes residual style information. Additionally, we deduplicate the same discrete token and employ a differentiable duration predictor to re-predict the duration of each token, which can adapt the duration of the same linguistic content to different styles. The subjective and objective evaluation results demonstrate the effectiveness of our proposed system.

📄 PDF Abstract BibTeX arXiv:2309.09262

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Latent Diffusion Model Diffusion models applied to latent spaces, which are normally built with (Variational) Autoencoders.

Similar Papers 제목 키워드 기반

StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion

2021-07-21 · Yinghao Aaron Li, Ali Zare, Nima Mesgarani

We present an unsupervised non-parallel many-to-many voice conversion (VC) method using a generative adversarial network (GAN) called StarGAN v2. Using a combination of adversarial source classifier loss and perceptual l…

Generative Adversarial Networktext-to-speechText to SpeechVoice Conversion

Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion

2025-07-19 · Yu Zhang, Baotong Tian, Zhiyao Duan arxiv

Zero-shot online voice conversion (VC) holds significant promise for real-time communications and entertainment. However, current VC models struggle to preserve semantic fidelity under real-time constraints, deliver natu…

Voice Conversion

Singing Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders

2019-12-03 · Yin-Jyun Luo, Chin-Chen Hsu, Kat Agres, Dorien Herremans

We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages re…

DecoderVoice Conversion

Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder

2025-04-16 · Soobin Suh, Dabi Ahn, Heewoong Park, Jonghun Park

Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script wi…

DiversityVoice Conversion

Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations

2018-04-09 · Ju-chieh Chou, Cheng-chieh Yeh, Hung-Yi Lee, Lin-shan Lee

Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for ea…

DecoderVoice Conversion