paper-with-me

Papers

Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription

2023-09-27 · Frank Cwitkowitz, Kin Wai Cheuk, Woosung Choi, Marco A. Martínez-Ramírez, Keisuke Toyama, Wei-Hsiang Liao, Yuki Mitsufuji

In recent years, research on music transcription has focused mainly on architecture design and instrument-specific data acquisition. With the lack of availability of diverse datasets, progress is often limited to solo-instrument tasks such as piano transcription. Several works have explored multi-instrument transcription as a means to bolster the performance of models on low-resource tasks, but these methods face the same data availability issues. We propose Timbre-Trap, a novel framework which unifies music transcription and audio reconstruction by exploiting the strong separability between pitch and timbre. We train a single autoencoder to simultaneously estimate pitch salience and reconstruct complex spectral coefficients, selecting between either output during the decoding stage via a simple switch mechanism. In this way, the model learns to produce coefficients corresponding to timbre-less audio, which can be interpreted as pitch salience. We demonstrate that the framework leads to performance comparable to state-of-the-art instrument-agnostic transcription methods, while only requiring a small amount of annotated data.

📄 PDF Abstract BibTeX arXiv:2309.15717

Code (0)

등록된 구현이 없습니다.

Tasks

Music Transcription

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders

2019-06-19 · Yin-Jyun Luo, Kat Agres, Dorien Herremans

In this paper, we learn disentangled representations of timbre and pitch for musical instrument sounds. We adapt a framework based on variational autoencoders with Gaussian mixture latent distributions. Specifically, we …

Decoder

DisMix: Disentangling Mixtures of Musical Instruments for Source-level Pitch and Timbre Manipulation

2024-08-20 · Yin-Jyun Luo, Kin Wai Cheuk, Woosung Choi, Toshimitsu Uesaka 외

Existing work on pitch and timbre disentanglement has been mostly focused on single-instrument music audio, excluding the cases where multiple instruments are presented. To fill the gap, we propose DisMix, a generative f…

AttributeDisentanglement

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems

2026-05-10 · Leduo Chen, Junchuan Zhao, Shengchen Li arxiv

Timbre transfer aims to modify the timbral identity of a musical recording while preserving the original melody and rhythm. While single-instrument timbre transfer has made substantial progress, existing approaches to mu…

Style Transfer

Timbre transfer using image-to-image denoising diffusion implicit models

2023-07-10 · Luca Comanducci, Fabio Antonacci, Augusto Sarti

Timbre transfer techniques aim at converting the sound of a musical piece generated by one instrument into the same one as if it was played by another instrument, while maintaining as much as possible the content in term…

DenoisingImage Denoising

Timbre Classification of Musical Instruments with a Deep Learning Multi-Head Attention-Based Model

2021-07-13 · Carlos Hernandez-Olivan, Jose R. Beltran

The aim of this work is to define a model based on deep learning that is able to identify different instrument timbres with as few parameters as possible. For this purpose, we have worked with classical orchestral instru…