paper-with-me

Papers

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

2024-09-20 · Yu Zhang, Changhao Pan, Wenxiang Guo, RuiQi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, Lichao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and realistic music scores, and poor task suitability. To tackle these problems, we present GTSinger, a large global, multi-technique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks. Particularly, (1) we collect 80.59 hours of high-quality singing voices, forming the largest recorded singing dataset; (2) 20 professional singers across nine widely spoken languages offer diverse timbres and styles; (3) we provide controlled comparison and phoneme-level annotations of six commonly used singing techniques, helping technique modeling and control; (4) GTSinger offers realistic music scores, assisting real-world musical composition; (5) singing voices are accompanied by manual phoneme-to-audio alignments, global style labels, and 16.16 hours of paired speech for various singing tasks. Moreover, to facilitate the use of GTSinger, we conduct four benchmark experiments: technique-controllable singing voice synthesis, technique recognition, style transfer, and speech-to-singing conversion. The corpus and demos can be found at http://aaronz345.github.io/GTSingerDemo/. We provide the dataset and the code for processing data and conducting benchmarks at https://huggingface.co/datasets/AaronZ345/GTSinger and https://github.com/AaronZ345/GTSinger.

📄 PDF Abstract BibTeX arXiv:2409.13832

Code (1)

aaronz345/gtsinger 공식 구현 pytorch

Tasks

AllSinging Voice SynthesisStyle TransferVocal technique classification

Similar Papers 제목 키워드 기반

SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling

2025-06-17 · Tawsif Ahmed, Andrej Radonjic, Gollam Rabby

We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song. To the best of our knowledge, there are no open-source high-quality dataset representing popular and well-known songs for generative mu…

Music CaptioningMusic ModelingSinging Voice Synthesis

WeSinger: Data-augmented Singing Voice Synthesis with Auxiliary Losses

2022-03-21 · Zewang Zhang, Yibin Zheng, Xinhui Li, Li Lu

In this paper, we develop a new multi-singer Chinese neural singing voice synthesis (SVS) system named WeSinger. To improve the accuracy and naturalness of synthesized singing voice, we design several specifical modules …

Data AugmentationDecoderRhythmSinging Voice Synthesis

M4Singer: a Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

2022-12-29 · NIPS 2022 12 · Lichao Zhang, RuiQi Li, Shoutong Wang, Liqun Deng 외

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Mult…

Music TranscriptionSinging Voice SynthesisVoice Conversion

Investigation of Singing Voice Separation for Singing Voice Detection in Polyphonic Music

2020-04-08 · Yifu Sun, xulong Zhang, Yi Yu, Xi Chen 외

Singing voice detection (SVD), to recognize vocal parts in the song, is an essential task in music information retrieval (MIR). The task remains challenging since singing voice varies and intertwines with the accompanime…

Information RetrievalMelody ExtractionMusic Information RetrievalRetrieval

Automatic Estimation of Singing Voice Musical Dynamics

2024-10-27 · Jyoti Narang, Nazif Can Tamer, Viviana De La Vega, Xavier Serra

Musical dynamics form a core part of expressive singing voice performances. However, automatic analysis of musical dynamics for singing voice has received limited attention partly due to the scarcity of suitable datasets…