ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference speed to exchange with high-quality sample generation limits its application scenarios. In order to obtain high quality synthetic singing voice more efficiently, we propose a singing voice synthesis method based on the consistency model, ConSinger, to achieve high-fidelity singing voice synthesis with minimal steps. The model is trained by applying consistency constraint and the generation quality is greatly improved at the expense of a small amount of inference speed. Our experiments show that ConSinger is highly competitive with the baseline model in terms of generation speed and quality. Audio samples are available at https://keylxiao.github.io/consinger.
Code (0)
등록된 구현이 없습니다.
Tasks
Singing Voice SynthesisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis
High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…
Singing Voice SynthesisVocal Bursts Intensity PredictionSingGAN: Generative Adversarial Network For High-Fidelity Singing Voice Generation
Deep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts…
Generative Adversarial NetworkGPUSinging Voice SynthesisSpeech Synthesis+3Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus
High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not me…
Audio GenerationSinging Voice SynthesisText-To-Speech SynthesisReal-Time and Accurate: Zero-shot High-Fidelity Singing Voice Conversion with Multi-Condition Flow Synthesis
Singing voice conversion is to convert the source singing voice into the target singing voice except for the content. Currently, flow-based models can complete the task of voice conversion, but they struggle to effective…
AttributeDecoderVoice ConversionHiFi-WaveGAN: Generative Adversarial Network with Auxiliary Spectrogram-Phase Loss for High-Fidelity Singing Voice Generation
Entertainment-oriented singing voice synthesis (SVS) requires a vocoder to generate high-fidelity (e.g. 48kHz) audio. However, most text-to-speech (TTS) vocoders cannot reconstruct the waveform well in this scenario. In …
Generative Adversarial NetworkSinging Voice Synthesistext-to-speechText to Speech