Boundary Regression for Leitmotif Detection in Music Audio
Leitmotifs are musical phrases that are reprised in various forms throughout a piece. Due to diverse variations and instrumentation, detecting the occurrence of leitmotifs from audio recordings is a highly challenging task. Leitmotif detection may be handled as a subcategory of audio event detection, where leitmotif activity is predicted at the frame level. However, as leitmotifs embody distinct, coherent musical structures, a more holistic approach akin to bounding box regression in visual object detection can be helpful. This method captures the entirety of a motif rather than fragmenting it into individual frames, thereby preserving its musical integrity and producing more useful predictions. We present our experimental results on tackling leitmotif detection as a boundary regression task.
Code (0)
등록된 구현이 없습니다.
Tasks
Event Detectionobject-detectionObject DetectionregressionSimilar Papers 제목 키워드 기반
Discovering Leitmotifs in Multidimensional Time Series
A leitmotif is a recurring theme in literature, movies or music that carries symbolic significance for the piece it is contained in. When this piece can be represented as a multi-dimensional time series (MDTS), such as a…
Time SeriesBarwise Section Boundary Detection in Symbolic Music Using Convolutional Neural Networks
Current methods for Music Structure Analysis (MSA) focus primarily on audio data. While symbolic music can be synthesized into audio and analyzed using existing MSA techniques, such an approach does not exploit symbolic …
Boundary DetectionEDMFormer: Genre-Specific Self-Supervised Learning for Music Structure Segmentation
Music structure segmentation is a key task in audio analysis, but existing models perform poorly on Electronic Dance Music (EDM). This problem exists because most approaches rely on lyrical or harmonic similarity, which …
Self-Supervised LearningBoundary DetectionEvaluating Multimodal Large Language Models on Core Music Perception Tasks
Multimodal Large Language Models (LLMs) claim "musical understanding" via evaluations that conflate listening with score reading. We benchmark three SOTA LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni) across t…
Audio Content Analysis
Preprint for a book chapter introducing Audio Content Analysis. With a focus on Music Information Retrieval systems, this chapter defines musical audio content, introduces the general process of audio content analysis, a…
Emotion RecognitionGeneral ClassificationGenre classificationInformation Retrieval+7