paper-with-me

홈 › Papers

Modeling Singing F0 With Neural Network Driven Transition-Sustain Models

2018-03-11

This study focuses on generating fundamental frequency (F0) curves of singing voice from musical scores stored in a midi-like notation. Current statistical parametric approaches to singing F0 modeling meet difficulties in reproducing vibratos and the temporal details at note boundaries due to the oversmoothing tendency of statistical models. This paper presents a neural network based solution that models a pair of neighboring notes at a time (the transition model) and uses a separate network for generating vibratos (the sustain model). Predictions from the two models are combined by summation after proper enveloping to enforce continuity. In the training phase, mild misalignment between the scores and the target F0 is addressed by back-propagating the gradients to the networks' inputs. Subjective listening tests on the NITech singing database show that transition-sustain models are able to generate F0 trajectories close to the original performance.

📄 PDF Abstract BibTeX arXiv:1803.04030

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

2025-05-20 · Yu Zhang, Wenxiang Guo, Changhao Pan, Dongyu Yao 외

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary a…

Contrastive LearningSinging Voice SynthesisStyle Transfer

SingingHead: A Large-scale 4D Dataset for Singing Head Animation

2023-12-07 · Sijing Wu, Yunhao Li, Weitian Zhang, Jun Jia 외

Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, i…

Portrait AnimationRhythm

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

2025-09-02 · Zikai Huang, Yihan Zhou, Xuemiao Xu, Cheng Xu 외 arxiv

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric…

Singing-Tacotron: Global duration control attention and dynamic filter for End-to-end singing voice synthesis

2022-02-16 · Tao Wang, Ruibo Fu, Jiangyan Yi, JianHua Tao 외

End-to-end singing voice synthesis (SVS) is attractive due to the avoidance of pre-aligned data. However, the auto learned alignment of singing voice with lyrics is difficult to match the duration information in musical …

Singing Voice Synthesis

InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

2026-09-04 · Yihan Zhou, Zikai Huang, Yuyang Yu, Xuemiao Xu 외 arxiv

We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. Unlike solo singing, duet performance requires each singer to balance individual expressiveness with intermitte…