paper-with-me

홈 › Papers

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

2026-08-06 · Eoin Cummins, Zhongyi Huang, Alexandre D'Hooge, Zhuoro Mo, Yaolong Ju arxiv

Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with \texttt{**kern} score encodings for 9,468 clips originating from 6,066 unique songs, the first of its kind to facilitate A2S research for popular music. Additionally, we improve on existing A2S approaches by using data augmentation and MuQ, a pretrained feature-extraction model for music audio, to enhance generalisation abilities and extract meaningful audio features. Results show that the proposed A2S model achieves 4.98\% symbol error rate (SER) on the Quartets collection for classical music, which significantly outperforms the 15.3\% SER from the existing state-of-the-art \cite{alfaro-contrerasTransformer2024}. Additionally, our model achieves 20.92\% SER on the SheetSage-A2S dataset for popular music, serving as a strong benchmark for future research. The dataset, model, and code are made publicly available at: https://github.com/Multimodal-Music-Research-Lab/SheetSage2Kern_model.

📄 PDF Abstract BibTeX arXiv:2608.06165

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding

2024-05-22 · Wei Zeng, Xian He, Ye Wang

Piano audio-to-score transcription (A2S) is an important yet underexplored task with extensive applications for music composition, practice, and analysis. However, existing end-to-end piano A2S systems faced difficulties…

DecoderMulti-Task LearningMusic Transcription

RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection

2025-07-16 · Sungkyun Chang, Simon Dixon, Emmanouil Benetos arxiv

This study introduces RUMAA, a transformer-based framework for music performance analysis that unifies score-to-performance alignment, score-informed transcription, and mistake detection in a near end-to-end manner. Unli…

Automatic Music Transcription using Convolutional Neural Networks and Constant-Q transform

2025-05-07 · Yohannis Telila, Tommaso Cucinotta, Davide Bacciu

Automatic music transcription (AMT) is the problem of analyzing an audio recording of a musical piece and detecting notes that are being played. AMT is a challenging problem, particularly when it comes to polyphonic musi…

Music Transcription

Music Transcription with (Almost) No Supervision

2026-05-22 · Saebyeol Shin, Chao Wan, Zhenzhen Liu, Justin Lovelace 외 arxiv

Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions. Meanwhile, vast quantities of unpaired a…

Music Transcription

A holistic approach to polyphonic music transcription with neural networks

2019-10-26 · Miguel A. Román, Antonio Pertusa, Jorge Calvo-Zaragoza

We present a framework based on neural networks to extract music scores directly from polyphonic audio in an end-to-end fashion. Most previous Automatic Music Transcription (AMT) methods seek a piano-roll representation …

Beat TrackingMusic TranscriptionQuantizationRhythm