paper-with-me

Papers

ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps

2024-10-20 · Yulin Song, Guorui Sang, Jing Yu, Chuangbai Xiao

Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference speed to exchange with high-quality sample generation limits its application scenarios. In order to obtain high quality synthetic singing voice more efficiently, we propose a singing voice synthesis method based on the consistency model, ConSinger, to achieve high-fidelity singing voice synthesis with minimal steps. The model is trained by applying consistency constraint and the generation quality is greatly improved at the expense of a small amount of inference speed. Our experiments show that ConSinger is highly competitive with the baseline model in terms of generation speed and quality. Audio samples are available at https://keylxiao.github.io/consinger.

📄 PDF Abstract BibTeX arXiv:2410.15342

Code (0)

등록된 구현이 없습니다.

Tasks

Singing Voice Synthesis

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

2020-09-03 · Jiawei Chen, Xu Tan, Jian Luan, Tao Qin 외

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…

Singing Voice SynthesisVocal Bursts Intensity Prediction

SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice Generation

2021-10-14 · Rongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren 외

Deep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts…

Generative Adversarial NetworkGPUSinging Voice SynthesisSpeech Synthesis+3

Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus

2021-12-20 · MM '21: Proceedings of the 29th ACM International Conference on Multimedia 2021 10 · Rongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu 외

High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not me…

Audio GenerationSinging Voice SynthesisText-To-Speech Synthesis

Real-Time and Accurate: Zero-shot High-Fidelity Singing Voice Conversion with Multi-Condition Flow Synthesis

2024-05-23 · Hui Li, Hongyu Wang, Zhijin Chen, Bohan Sun 외

Singing voice conversion is to convert the source singing voice into the target singing voice except for the content. Currently, flow-based models can complete the task of voice conversion, but they struggle to effective…

AttributeDecoderVoice Conversion

HiFi-WaveGAN: Generative Adversarial Network with Auxiliary Spectrogram-Phase Loss for High-Fidelity Singing Voice Generation

2022-10-23 · Chunhui Wang, Chang Zeng, Jun Chen, Xing He

Entertainment-oriented singing voice synthesis (SVS) requires a vocoder to generate high-fidelity (e.g. 48kHz) audio. However, most text-to-speech (TTS) vocoders cannot reconstruct the waveform well in this scenario. In …

Generative Adversarial NetworkSinging Voice Synthesistext-to-speechText to Speech