paper-with-me

홈 › Papers

Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence

2025-09-08 · Yerin Ryu, Inseop Shin, Chanwoo Kim arxiv

Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise control over attributes such as dynamics. We address this by focusing on dynamic control--temporal loudness variation essential for musical expressiveness--and explicitly condition the SVS model on energy sequences extracted from ground-truth spectrograms, reducing annotation costs and improving controllability. We also propose a phoneme-level energy sequence for user-friendly control. To the best of our knowledge, this is the first attempt enabling user-driven dynamics control in SVS. Experiments show our method achieves over 50% reduction in mean absolute error of energy sequences for phoneme-level inputs compared to baseline and energy-predictor models, without compromising synthesis quality.

📄 PDF Abstract BibTeX arXiv:2509.07038

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

2026-07-29 · Wei-Jaw Lee, Hsuan-Yu Yeh, Ting-Yi Hu, Chih-Pin Tan 외 arxiv

Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components. The state-of-the-art model SongEcho utilizes $F_0$ sequences and voice…

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

2024-09-20 · Yu Zhang, Changhao Pan, Wenxiang Guo, RuiQi Li 외

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited div…

AllSinging Voice SynthesisStyle TransferVocal technique classification

RMSSinger: Realistic-Music-Score based Singing Voice Synthesis

2023-05-18 · Jinzheng He, Jinglin Liu, Zhenhui Ye, Rongjie Huang 외

We are interested in a challenging task, Realistic-Music-Score based Singing Voice Synthesis (RMS-SVS). RMS-SVS aims to generate high-quality singing voices given realistic music scores with different note types (grace, …

Singing Voice Synthesis

A Singing Voice Database in Basque for Statistical Singing Synthesis of Bertsolaritza

2016-05-01 · LREC 2016 5 · Xabier Sarasola, Eva Navas, David Tavarez, Daniel Erro 외

This paper describes the characteristics and structure of a Basque singing voice database of bertsolaritza. Bertsolaritza is a popular singing style from Basque Country sung exclusively in Basque that is improvised and a…

Singing Voice Synthesis

Rapping-Singing Voice Synthesis based on Phoneme-level Prosody Control

2021-11-17 · Konstantinos Markopoulos, Nikolaos Ellinas, Alexandra Vioni, Myrsini Christidou 외

In this paper, a text-to-rapping/singing system is introduced, which can be adapted to any speaker's voice. It utilizes a Tacotron-based multispeaker acoustic model trained on read-only speech data and which provides pro…

Singing Voice Synthesisvalid