paper-with-me

홈 › Papers

PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective

2023-09-05 · Alain Riou, Stefan Lattner, Gaëtan Hadjeres, Geoffroy Peeters

In this paper, we address the problem of pitch estimation using Self Supervised Learning (SSL). The SSL paradigm we use is equivariance to pitch transposition, which enables our model to accurately perform pitch estimation on monophonic audio after being trained only on a small unlabeled dataset. We use a lightweight ($<$ 30k parameters) Siamese neural network that takes as inputs two different pitch-shifted versions of the same audio represented by its Constant-Q Transform. To prevent the model from collapsing in an encoder-only setting, we propose a novel class-based transposition-equivariant objective which captures pitch information. Furthermore, we design the architecture of our network to be transposition-preserving by introducing learnable Toeplitz matrices. We evaluate our model for the two tasks of singing voice and musical instrument pitch estimation and show that our model is able to generalize across tasks and datasets while being lightweight, hence remaining compatible with low-resource devices and suitable for real-time applications. In particular, our results surpass self-supervised baselines and narrow the performance gap between self-supervised and supervised methods for pitch estimation.

📄 PDF Abstract BibTeX arXiv:2309.02265

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective

2025-08-02 · Alain Riou, Bernardo Torres, Ben Hayes, Stefan Lattner 외 arxiv

In this paper, we introduce PESTO, a self-supervised learning approach for single-pitch estimation using a Siamese architecture. Our model processes individual frames of a Variable-$Q$ Transform (VQT) and predicts pitch …

Self-Supervised Learning

Learning Transposition-Invariant Interval Features from Symbolic Music and Audio

2018-06-21 · Stefan Lattner, Maarten Grachten, Gerhard Widmer

Many music theoretical constructs (such as scale types, modes, cadences, and chord types) are defined in terms of pitch intervals---relative distances between pitches. Therefore, when computer models are employed in musi…

Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music

2026-01-16 · Venkat Suprabath Bitra, Homayoon Beigi arxiv

Reliable fundamental frequency (F 0) and voicing estimation is essential for neural synthesis, yet many pitch extractors depend on large labeled corpora and degrade under realistic recording artifacts. We propose a light…

pyPESTO: A modular and scalable tool for parameter estimation for dynamic models

2023-05-02 · Yannik Schälte, Fabian Fröhlich, Paul J. Jost, Jakob Vanhoefer 외

Mechanistic models are important tools to describe and understand biological processes. However, they typically rely on unknown parameters, the estimation of which can be challenging for large and complex systems. We pre…

parameter estimationUncertainty Quantification

Toward Fully Self-Supervised Multi-Pitch Estimation

2024-02-23 · Frank Cwitkowitz, Zhiyao Duan

Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstr…

Self-Supervised Learning