paper-with-me

홈 › Papers

Transcoda: End-to-End Zero-Shot Optical Music Recognition via Data-Centric Synthetic Training

2026-05-11 · Daniel Dratschuk, Paul Swoboda arxiv

Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of large-scale, annotated datasets of real scans. This forces models to rely on either few-shot transfer or synthetic training pipelines that remain overly simplistic. A secondary challenge is encoding non-uniqueness: in the popular Humdrum kern format for transcribing music, multiple different text encodings can render into the same visual sheet music. This one-to-many mapping creates a harder learning task and introduces high uncertainty during decoding. We propose Transcoda, an OMR system built on (i) an advanced synthetic data generation pipeline, (ii) a normalization of the kern encoding to enforce a unique normal form and (iii) grammar-based decoding to ensure the syntactic correctness of the output. This approach allows us to train a compact 59M-parameter model in just 6 hours on a single GPU that outperforms billion-parameter baselines. Transcoda achieves the best score among state of the art baselines on a newly curated benchmark of synthetically rendered scores at 18.46% OMR-NED (compared to 43.91% for the next-best system, Legato) and reduces the error rate on historical Polish scans to 63.97% OMR-NED (down from 80.16% for SMT++).

📄 PDF Abstract BibTeX arXiv:2605.10835

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data Generation

Similar Papers 제목 키워드 기반

End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music

2024-05-20 · Antonio Ríos-Vila, Jorge Calvo-Zaragoza, David Rizo, Thierry Paquet

Optical Music Recognition (OMR) has made significant progress since its inception, with various approaches now capable of accurately transcribing music scores into digital formats. Despite these advancements, most so-cal…

Synthetic Data Generation

Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction

2024-10-15 · Renhang Liu, Abhinaba Roy, Dorien Herremans

In this work, we present a novel method for music emotion recognition that leverages Large Language Model (LLM) embeddings for label alignment across multiple datasets and zero-shot prediction on novel categories. First,…

Emotion RecognitionLanguage ModelingLanguage ModellingLarge Language Model+1

I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition

2024-07-25 · Yannis Vasilakis, Rachel Bittner, Johan Pauwels

Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels. These systems enable new approaches for class…

Instrument RecognitionRetrievalzero-shot-classificationZero-Shot Learning

TrOMR:Transformer-Based Polyphonic Optical Music Recognition

2023-08-18 · Yixuan Li, Huaping Liu, Qiang Jin, Miaomiao Cai 외

Optical Music Recognition (OMR) is an important technology in music and has been researched for a long time. Previous approaches for OMR are usually based on CNN for image understanding and RNN for music symbol classific…

Proceedings of the 4th International Workshop on Reading Music Systems

2022-11-23 · Jorge Calvo-Zaragoza, Alexander Pacha, Elona Shatri

The International Workshop on Reading Music Systems (WoRMS) is a workshop that tries to connect researchers who develop systems for reading music, such as in the field of Optical Music Recognition, with other researchers…

Information RetrievalMusic Information RetrievalRetrieval