paper-with-me

홈 › Papers

Towards end-to-end F0 voice conversion based on Dual-GAN with convolutional wavelet kernels

2021-04-15 · Clément Le Moine Veillon, Nicolas Obin, Axel Roebel

This paper presents a end-to-end framework for the F0 transformation in the context of expressive voice conversion. A single neural network is proposed, in which a first module is used to learn F0 representation over different temporal scales and a second adversarial module is used to learn the transformation from one emotion to another. The first module is composed of a convolution layer with wavelet kernels so that the various temporal scales of F0 variations can be efficiently encoded. The single decomposition/transformation network allows to learn in a end-to-end manner the F0 decomposition that are optimal with respect to the transformation, directly from the raw F0 signal.

📄 PDF Abstract BibTeX arXiv:2104.07283

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Spectrum and Prosody Conversion for Cross-lingual Voice Conversion with CycleGAN

2020-08-11 · Zongyang Du, Kun Zhou, Berrak Sisman, Haizhou Li

Cross-lingual voice conversion aims to change source speaker's voice to sound like that of target speaker, when source and target speakers speak different languages. It relies on non-parallel training data from two diffe…

Voice Conversion

VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion

2025-05-27 · Joon-Seung Choi, Dong-Min Byun, Hyung-Seok Oh, Seong-Whan Lee

Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling v…

Voice Conversion

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

2020-05-13 · Kun Zhou, Berrak Sisman, Mingyang Zhang, Haizhou Li

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried …

DecoderVoice Conversion

Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

2020-02-01 · Kun Zhou, Berrak Sisman, Haizhou Li

Emotional voice conversion aims to convert the spectrum and prosody to change the emotional patterns of speech, while preserving the speaker identity and linguistic content. Many studies require parallel speech data betw…

Voice Conversion

DualVC: Dual-mode Voice Conversion using Intra-model Knowledge Distillation and Hybrid Predictive Coding

2023-05-21 · Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu, Jixun Yao 외

Voice conversion is an increasingly popular technology, and the growing number of real-time applications requires models with streaming conversion capabilities. Unlike typical (non-streaming) voice conversion, which can …

Data AugmentationDecoderKnowledge DistillationVoice Conversion