paper-with-me

Papers

Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion

2024-06-04 · RuiQi Li, Rongjie Huang, Yongqi Wang, Zhiqing Hong, Zhou Zhao

Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal quality of generated outputs, presenting significant hurdles in STS research. This paper presents SVPT, an STS approach boosted by a self-supervised singing voice pre-training model. We leverage spoken language model techniques to tackle the rhythm alignment problem and the in-context learning capability to achieve zero-shot conversion. We adopt discrete-unit random resampling and pitch corruption strategies, enabling training with unpaired singing data and thus mitigating the issue of data scarcity. SVPT also serves as an effective backbone for singing voice synthesis (SVS), offering insights into scaling up SVS models. Experimental results indicate that SVPT delivers notable improvements in both STS and SVS endeavors. Audio samples are available at https://speech2sing.github.io.

📄 PDF Abstract BibTeX arXiv:2406.02429

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningLanguage ModelingLanguage ModellingRhythmSinging Voice SynthesisSTSVoice Conversion

Similar Papers 제목 키워드 기반

Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations

2024-02-02 · Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis, Myrsini Christidou 외

In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resource pipeline that does not utilize any sin…

Singing Voice Synthesis

Singer Identity Representation Learning using Self-Supervised Techniques

2024-01-10 · International Society of Music Information Retrieval 2023 8 · Bernardo Torres, Stefan Lattner, Gaël Richard

Significant strides have been made in creating voice identity representations using speech data. However, the same level of progress has not been achieved for singing voices. To bridge this gap, we suggest a framework fo…

Domain GeneralizationRepresentation LearningSelf-Supervised LearningSpeaker Verification+1

Singing Beat Tracking With Self-supervised Front-end and Linear Transformers

2022-08-31 · Mojtaba Heydari, Zhiyao Duan

Tracking beats of singing voices without the presence of musical accompaniment can find many applications in music production, automatic song arrangement, and social media interaction. Its main challenge is the lack of s…

Beat Tracking

Toward Leveraging Pre-Trained Self-Supervised Frontends for Automatic Singing Voice Understanding Tasks: Three Case Studies

2023-06-22 · Yuya Yamamoto

Automatic singing voice understanding tasks, such as singer identification, singing voice transcription, and singing technique classification, benefit from data-driven approaches that utilize deep learning techniques. Th…

DiversityMusic ClassificationSelf-Supervised LearningSinger Identification

MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance

2024-06-10 · Semin Kim, Myeonghun Jeong, Hyeonseung Lee, Minchan Kim 외

In this paper, we propose MakeSinger, a semi-supervised training method for singing voice synthesis (SVS) via classifier-free diffusion guidance. The challenge in SVS lies in the costly process of gathering aligned sets …

Singing Voice Synthesistext-to-speechText to Speech