paper-with-me

홈 › Papers

Multi-modal Conditional Bounding Box Regression for Music Score Following

2021-05-10 · Florian Henkel, Gerhard Widmer

This paper addresses the problem of sheet-image-based on-line audio-to-score alignment also known as score following. Drawing inspiration from object detection, a conditional neural network architecture is proposed that directly predicts x,y coordinates of the matching positions in a complete score sheet image at each point in time for a given musical performance. Experiments are conducted on a synthetic polyphonic piano benchmark dataset and the new method is compared to several existing approaches from the literature for sheet-image-based score following as well as an Optical Music Recognition baseline. The proposed approach achieves new state-of-the-art results and furthermore significantly improves the alignment performance on a set of real-world piano recordings by applying Impulse Responses as a data augmentation technique.

📄 PDF Abstract BibTeX arXiv:2105.04309

Code (1)

CPJKU/cyolo_score_following 공식 구현 pytorch

Tasks

Data Augmentationobject-detectionObject Detectionregression

Similar Papers 제목 키워드 기반

Detecting Noteheads in Handwritten Scores with ConvNets and Bounding Box Regression

2017-08-05 · Jan Hajič jr., Pavel Pecina

Noteheads are the interface between the written score and music. Each notehead on the page signifies one note to be played, and detecting noteheads is thus an unavoidable step for Optical Music Recognition. Noteheads are…

General Classificationregression

Boundary Regression for Leitmotif Detection in Music Audio

2025-03-11 · SiHun Lee, Dasaem Jeong

Leitmotifs are musical phrases that are reprised in various forms throughout a piece. Due to diverse variations and instrumentation, detecting the occurrence of leitmotifs from audio recordings is a highly challenging ta…

Event Detectionobject-detectionObject Detectionregression

Multi-Modality in Music: Predicting Emotion in Music from High-Level Audio Features and Lyrics

2023-02-26 · Tibor Krols, Yana Nikolova, Ninell Oldenburg

This paper aims to test whether a multi-modal approach for music emotion recognition (MER) performs better than a uni-modal one on high-level song features and lyrics. We use 11 song features retrieved from the Spotify A…

Emotion RecognitionMusic Emotion Recognitionregression

Multi-Modal Pedestrian Detection with Large Misalignment Based on Modal-Wise Regression and Multi-Modal IoU

2021-07-23 · Napat Wanchaitanawong, Masayuki Tanaka, Takashi Shibata, Masatoshi Okutomi

The combined use of multiple modalities enables accurate pedestrian detection under poor lighting conditions by using the high visibility areas from these modalities together. The vital assumption for the combination use…

Pedestrian Detectionregression

Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model

2023-11-02 · Jaeyong Kang, Soujanya Poria, Dorien Herremans

Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos. In this work, we develop a generative …

Music GenerationRhythm