paper-with-me

홈 › Papers

Deep Autotuner: A Data-Driven Approach to Natural-Sounding Pitch Correction for Singing Voice in Karaoke Performances

2019-02-03 · Sanna Wager, George Tzanetakis, Cheng-i Wang, Lijiang Guo, Aswin Sivaraman, Minje Kim

We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The proposed approach addresses the situation where no musical score of the vocals nor the accompaniment exists: It predicts the amount of correction from the relationship between the spectral contents of the vocal and accompaniment tracks. Hence, the pitch shift in cents suggested by the model can be used to make the voice sound in tune with the accompaniment. This approach differs from commercially used automatic pitch correction systems, where notes in the vocal tracks are shifted to be centered around notes in a user-defined score or mapped to the closest pitch among the twelve equal-tempered scale degrees. We train the model using a dataset of 4,702 amateur karaoke performances selected for good intonation. We present a Convolutional Gated Recurrent Unit (CGRU) model to accomplish this task. This method can be extended into unsupervised pitch correction of a vocal performance, popularly referred to as autotuning.

📄 PDF Abstract BibTeX arXiv:1902.00956

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Autotuner: a Pitch Correcting Network for Singing Performances

2020-02-12 · Sanna Wager, George Tzanetakis, Cheng-i Wang, Minje Kim

We introduce a data-driven approach to automatic pitch correction of solo singing performances. The proposed approach predicts note-wise pitch shifts from the relationship between the respective spectrograms of the singi…

Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch

2024-10-09 · Teodora Răgman, Adriana Stan

This paper focuses on adapting the functionalities of the FastPitch model to the Romanian language; extending the set of speakers from one to eighteen; synthesising speech using an anonymous identity; and replicating the…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Onsets and Frames: Dual-Objective Piano Transcription

2017-10-30 · Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts 외

We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset eve…

Music Transcription

Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis

2024-06-14 · Taewoo Kim, Choongsang Cho, Young Han Lee

In this paper, we present Period Singer, a novel end-to-end singing voice synthesis (SVS) model that utilizes variational inference for periodic and aperiodic components, aimed at producing natural-sounding waveforms. Re…

Singing Voice SynthesisVariational Inference

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2