paper-with-me

Papers

S2Cap: A Benchmark and a Baseline for Singing Style Captioning

2024-09-15 · Hyunjong Ok, Jaeho Lee

Singing voices contain much richer information than common voices, such as diverse vocal and acoustic characteristics. However, existing open-source audio-text datasets for singing voices capture only a limited set of attributes and lacks acoustic features, leading to limited utility towards downstream tasks, such as style captioning. To fill this gap, we formally consider the task of singing style captioning and introduce S2Cap, a singing voice dataset with comprehensive descriptions of diverse vocal, acoustic and demographic attributes. Based on this dataset, we develop a simple yet effective baseline algorithm for the singing style captioning. The algorithm utilizes two novel technical components: CRESCENDO for mitigating misalignment between pretrained unimodal models, and demixing supervision to regularize the model to focus on the singing voice. Despite its simplicity, the proposed method outperforms state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2409.09866

Code (0)

등록된 구현이 없습니다.

Tasks

Singing Voice Synthesis

Methods 이 논문이 사용한 방법론

Focus 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

2023-12-17 · Yu Zhang, Rongjie Huang, RuiQi Li, Jinzheng He 외

Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from ref…

QuantizationSinging Voice SynthesisStyle Transfer

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

2024-09-24 · Yu Zhang, Ziyue Jiang, RuiQi Li, Changhao Pan 외

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunc…

ClusteringLanguage ModellingQuantizationRhythm+2

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

2024-09-20 · Yu Zhang, Changhao Pan, Wenxiang Guo, RuiQi Li 외

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited div…

AllSinging Voice SynthesisStyle TransferVocal technique classification

M4Singer: a Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus

2022-12-29 · NIPS 2022 12 · Lichao Zhang, RuiQi Li, Shoutong Wang, Liqun Deng 외

The lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Mult…

Music TranscriptionSinging Voice SynthesisVoice Conversion

Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control

2026-06-15 · Joon-Seung Choi, Dong-Min Byun, Seong-Whan Lee arxiv

Singing style is a crucial aspect of a natural and expressive singing voice. Singers utilize singing styles to convey the feeling or emotion of the songs. Several works have been proposed to control singing style for mak…

Voice Conversion