paper-with-me

홈 › Papers

Music Transcription with (Almost) No Supervision

2026-05-22 · Saebyeol Shin, Chao Wan, Zhenzhen Liu, Justin Lovelace, Daniel C. Lin, Kilian Q. Weinberger, John Thickstun arxiv

Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions. Meanwhile, vast quantities of unpaired audio recordings and symbolic scores are freely available but have gone unused. We adopt a cycle-consistent translation framework in which a small amount of paired data acts as a minimal anchor, unlocking the full potential of the unpaired pool. We find that: unpaired data yields surprisingly large gains, especially under limited supervision; unpaired audio contributes more than unpaired scores; incorporating unlabeled audio from a new instrument during training improves transcription for that instrument without any paired supervision. Together, these results suggest that scaling unpaired data offers a practical path toward high-quality transcription for instruments where labeled data remains scarce.

📄 PDF Abstract BibTeX arXiv:2605.24193

Code (0)

등록된 구현이 없습니다.

Tasks

Music Transcription

Similar Papers 제목 키워드 기반

Unaligned Supervision For Automatic Music Transcription in The Wild

2022-04-28 · Ben Maman, Amit H. Bermano

Multi-instrument Automatic Music Transcription (AMT), or the decoding of a musical recording into semantic musical content, is one of the holy grails of Music Information Retrieval. Current AMT approaches are restricted …

Information RetrievalMusic Information RetrievalMusic TranscriptionRetrieval

Invariances and Data Augmentation for Supervised Music Transcription

2017-11-13 · John Thickstun, Zaid Harchaoui, Dean Foster, Sham M. Kakade

This paper explores a variety of models for frame-based music transcription, with an emphasis on the methods needed to reach state-of-the-art on human recordings. The translation-invariant network discussed in this paper…

Data AugmentationMusic TranscriptionTranslation

Count The Notes: Histogram-Based Supervision for Automatic Music Transcription

2025-11-18 · Jonathan Yaffe, Ben Maman, Meinard Müller, Amit H. Bermano arxiv

Automatic Music Transcription (AMT) converts audio recordings into symbolic musical representations. Training deep neural networks (DNNs) for AMT typically requires strongly aligned training pairs with precise frame-leve…

Music Transcription

Transcription Is All You Need: Learning to Separate Musical Mixtures with Score as Supervision

2020-10-22 · Yun-Ning Hung, Gordon Wichern, Jonathan Le Roux

Most music source separation systems require large collections of isolated sources for training, which can be difficult to obtain. In this work, we use musical scores, which are comparatively easy to obtain, as a weak la…

AllMusic Source Separation

Music transcription modelling and composition using deep learning

2016-04-29 · Bob L. Sturm, João Felipe Santos, Oded Ben-Tal, Iryna Korshunova

We apply deep learning methods, specifically long short-term memory (LSTM) networks, to music transcription modelling and composition. We build and train LSTM networks using approximately 23,000 music transcriptions expr…

Deep LearningDescriptiveMusic Transcription