paper-with-me

홈 › Papers

FastSVC: Fast Cross-Domain Singing Voice Conversion with Feature-wise Linear Modulation

2020-11-11 · Songxiang Liu, Yuewen Cao, Na Hu, Dan Su, Helen Meng

This paper presents FastSVC, a light-weight cross-domain singing voice conversion (SVC) system, which can achieve high conversion performance, with inference speed 4x faster than real-time on CPUs. FastSVC uses Conformer-based phoneme recognizer to extract singer-agnostic linguistic features from singing signals. A feature-wise linear modulation based generator is used to synthesize waveform directly from linguistic features, leveraging information from sine-excitation signals and loudness features. The waveform generator can be trained conveniently using a multi-resolution spectral loss and an adversarial loss. Experimental results show that the proposed FastSVC system, compared with a computationally heavy baseline system, can achieve comparable conversion performance in some scenarios and significantly better conversion performance in other scenarios. Moreover, the proposed FastSVC system achieves desirable cross-lingual singing conversion performance. The inference speed of the FastSVC system is 3x and 70x faster than the baseline system on GPUs and CPUs, respectively.

📄 PDF Abstract BibTeX arXiv:2011.05731

Code (2)

lesterphillip/svcc23_fastsvc pytorch
liusongxiang/ppg-vc pytorch

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus

2021-12-20 · MM '21: Proceedings of the 29th ACM International Conference on Multimedia 2021 10 · Rongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu 외

High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not me…

Audio GenerationSinging Voice SynthesisText-To-Speech Synthesis

PolySinger: Singing-Voice to Singing-Voice Translation from English to Japanese

2024-07-19 · Silas Antonisen, Iván López-Espejo

The speech domain prevails in the spotlight for several natural language processing (NLP) tasks while the singing domain remains less explored. The culmination of NLP is the speech-to-speech translation (S2ST) task, refe…

Singing Voice SynthesisSpeech-to-Speech TranslationTranslation

Fast and High-Quality Singing Voice Synthesis System based on Convolutional Neural Networks

2019-10-24 · Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura 외

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the n…

Singing Voice Synthesis

A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023

2023-10-08 · Ryuichi Yamamoto, Reo Yoneyama, Lester Phillip Violeta, Wen-Chin Huang 외

This paper presents our systems (denoted as T13) for the singing voice conversion challenge (SVCC) 2023. For both in-domain and cross-domain English singing voice conversion (SVC) tasks (Task 1 and Task 2), we adopt a re…

Self-Supervised LearningTask 2Voice Conversion

StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

2023-12-17 · Yu Zhang, Rongjie Huang, RuiQi Li, Jinzheng He 외

Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from ref…

QuantizationSinging Voice SynthesisStyle Transfer