paper-with-me

홈 › Papers

Cross-domain Neural Pitch and Periodicity Estimation

2023-01-28 · Max Morrison, Caedon Hsieh, Nathan Pruyne, Bryan Pardo

Pitch is a foundational aspect of our perception of audio signals. Pitch contours are commonly used to analyze speech and music signals and as input features for many audio tasks, including music transcription, singing voice synthesis, and prosody editing. In this paper, we describe a set of techniques for improving the accuracy of widely-used neural pitch and periodicity estimators to achieve state-of-the-art performance on both speech and music. We also introduce a novel entropy-based method for extracting periodicity and per-frame voiced-unvoiced classifications from statistical inference-based pitch estimators (e.g., neural networks), and show how to train a neural pitch estimator to simultaneously handle both speech and music data (i.e., cross-domain estimation) without performance degradation. Our estimator implementations run 11.2x faster than real-time on a Intel i9-9820X 10-core 3.30 GHz CPU$\unicode{x2014}$approaching the speed of state-of-the-art DSP-based pitch estimators$\unicode{x2014}$or 408x faster than real-time on a NVIDIA GeForce RTX 3090 GPU. We release all of our code and models as Pitch-Estimating Neural Networks (penn), an open-source, pip-installable Python module for training, evaluating, and performing inference with pitch- and periodicity-estimating neural networks. The code for penn is available at https://github.com/interactiveaudiolab/penn.

📄 PDF Abstract BibTeX arXiv:2301.12258

Code (1)

interactiveaudiolab/penn 공식 구현 pytorch

Tasks

CPUGPUMusic TranscriptionSinging Voice Synthesis

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion

2023-06-16 · Woo-Jin Chung, Doyeon Kim, Soo-Whan Chung, Hong-Goo Kang

We introduce Multi-level feature Fusion-based Periodicity Analysis Model (MF-PAM), a novel deep learning-based pitch estimation model that accurately estimates pitch trajectory in noisy and reverberant acoustic environme…

Audio Signal Processing

Pitch and timbre discrimination at wave-to-spike transition in the cochlea

2017-11-15

A new definition of musical pitch is proposed. A Finite-Difference Time Domain (FDTM) model of the cochlea is used to calculate spike trains caused by tone complexes and by a recorded classical guitar tone. All harmonic …

Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis

2022-10-28 · Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima 외

Several fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately). However, they often generate …

DecoderDiversityEmotional Speech SynthesisSpeech Synthesis+3

Chunked Autoregressive GAN for Conditional Waveform Synthesis

2021-10-19 · ICLR 2022 4 · Max Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman 외

Conditional waveform synthesis models learn a distribution of audio waveforms given conditioning such as text, mel-spectrograms, or MIDI. These systems employ deep generative models that model the waveform via either seq…

Inductive Bias

hf0: A hybrid pitch extraction method for multimodal voice

2019-04-22 · Pradeep Rengaswamy, Gurunath Reddy M, Krothapalli Sreenivasa Rao

Pitch or fundamental frequency (f0) extraction is a fundamental problem studied extensively for its potential applications in speech and clinical applications. In literature, explicit mode specific (modal speech or singi…

Deep Learning