paper-with-me

홈 › Papers

Deep Autotuner: a Pitch Correcting Network for Singing Performances

2020-02-12 · Sanna Wager, George Tzanetakis, Cheng-i Wang, Minje Kim

We introduce a data-driven approach to automatic pitch correction of solo singing performances. The proposed approach predicts note-wise pitch shifts from the relationship between the respective spectrograms of the singing and accompaniment. This approach differs from commercial systems, where vocal track notes are usually shifted to be centered around pitches in a user-defined score, or mapped to the closest pitch among the twelve equal-tempered scale degrees. The proposed system treats pitch as a continuous value rather than relying on a set of discretized notes found in musical scores, thus allowing for improvisation and harmonization in the singing performance. We train our neural network model using a dataset of 4,702 amateur karaoke performances selected for good intonation. Our model is trained on both incorrect intonation, for which it learns a correction, and intentional pitch variation, which it learns to preserve. The proposed deep neural network with gated recurrent units on top of convolutional layers shows promising performance on the real-world score-free singing pitch correction task of autotuning.

📄 PDF Abstract BibTeX arXiv:2002.05511

Code (1)

sannawag/autotuner 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Deep Autotuner: A Data-Driven Approach to Natural-Sounding Pitch Correction for Singing Voice in Karaoke Performances

2019-02-03 · Sanna Wager, George Tzanetakis, Cheng-i Wang, Lijiang Guo 외

We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The proposed approach addresses the situation wh…

Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking

2019-02-09 · Hiroki Tamaru, Yuki Saito, Shinnosuke Takamichi, Tomoki Koriyama 외

This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of …

Singing Voice Synthesis

PitchNet: Unsupervised Singing Voice Conversion with Pitch Adversarial Network

2019-12-04 · Chengqi Deng, Chengzhu Yu, Heng Lu, Chao Weng 외

Singing voice conversion is to convert a singer's voice to another one's voice without changing singing content. Recent work shows that unsupervised singing voice conversion can be achieved with an autoencoder-based appr…

DecoderMusic GenerationTranslationVoice Conversion

Human Voice Pitch Estimation: A Convolutional Network with Auto-Labeled and Synthetic Data

2023-08-14 · Jeremy Cochoy

In the domain of music and sound processing, pitch extraction plays a pivotal role. Our research presents a specialized convolutional neural network designed for pitch extraction, particularly from the human singing voic…

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

2025-09-24 · Junchuan Zhao, Wei Zeng, Tianle Lyu, Ye Wang arxiv

Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled z…

Contrastive LearningSpeech Synthesis