paper-with-me

홈 › Papers

DoReMi: First glance at a universal OMR dataset

2021-07-16 · Elona Shatri, György Fazekas

The main challenges of Optical Music Recognition (OMR) come from the nature of written music, its complexity and the difficulty of finding an appropriate data representation. This paper provides a first look at DoReMi, an OMR dataset that addresses these challenges, and a baseline object detection model to assess its utility. Researchers often approach OMR following a set of small stages, given that existing data often do not satisfy broader research. We examine the possibility of changing this tendency by presenting more metadata. Our approach complements existing research; hence DoReMi allows harmonisation with two existing datasets, DeepScores and MUSCIMA++. DoReMi was generated using a music notation software and includes over 6400 printed sheet music images with accompanying metadata useful in OMR research. Our dataset provides OMR metadata, MIDI, MEI, MusicXML and PNG files, each aiding a different stage of OMR. We obtain 64% mean average precision (mAP) in object detection using half of the data. Further work includes re-iterating through the creation process to satisfy custom OMR models. While we do not assume to have solved the main challenges in OMR, this dataset opens a new course of discussions that would ultimately aid that goal.

📄 PDF Abstract BibTeX arXiv:2107.07786

Code (1)

apacha/OMR-Datasets

Tasks

object-detectionObject Detection

Methods 이 논문이 사용한 방법론

MEI MEI introduces the *multi-partition embedding interaction* technique with block term tensor format to systematically address the efficiency--expressiveness trade-off in…

Similar Papers 제목 키워드 기반

DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

2023-05-17 · NeurIPS 2023 11 · Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 외

The mixture proportions of pretraining data domains (e.g., Wikipedia, books, web text) greatly affect language model (LM) performance. In this paper, we propose Domain Reweighting with Minimax Optimization (DoReMi), whic…

Language ModelingLanguage Modelling

DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts

2025-11-14 · Mingwei Xing, Xinliang Wang, Yifeng Shi arxiv

Constructing a unified 3D scene understanding model has long been hindered by the significant topological discrepancies across different sensor modalities. While applying the Mixture-of-Experts (MoE) architecture is an e…

Scene Understanding

HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets

2025-03-16 · Lulu Ban, Tao Zhu, Xiangqing Lu, Qi Qiu 외

Cross-dataset Human Activity Recognition (HAR) suffers from limited model generalization, hindering its practical deployment. To address this critical challenge, inspired by the success of DoReMi in Large Language Models…

Activity RecognitionHuman Activity Recognition

DOREMI: Optimizing Long Tail Predictions in Document-Level Relation Extraction

2026-01-16 · Laura Menotti, Stefano Marchesin, Gianmaria Silvello arxiv

Document-Level Relation Extraction (DocRE) presents significant challenges due to its reliance on cross-sentence context and the long-tail distribution of relation types, where many relations have scarce training example…

Document-level Relation Extraction

DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment

2023-07-01 · Yanjiang Guo, Yen-Jen Wang, Lihan Zha, Jianyu Chen

Large language models (LLMs) encode a vast amount of semantic knowledge and possess remarkable understanding and reasoning capabilities. Previous work has explored how to ground LLMs in robotic tasks to generate feasible…

Language ModelingLanguage ModellingQuestion AnsweringVisual Question Answering (VQA)