paper-with-me

Papers

Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores

2025-11-24 · Congren Dai, Yue Yang, Krinos Li, Huichi Zhou, Shijie Liang, Bo Zhang, Enyang Liu, Ge Jin, Hongran An, Haosen Zhang, Peiyuan Jing, Kinhei Lee, Z henxuan Zhang, Xiaobing Li, Maosong Sun arxiv

Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains insufficiently examined. We introduce Musical Score Understanding Benchmark (MSU-Bench), a human-curated benchmark for score-level musical understanding across textual (ABC notation) and visual (PDF) modalities. MSU-Bench contains 1,800 generative question-answer pairs from works by Bach, Beethoven, Chopin, Debussy, and others, organised into four levels of increasing difficulty, ranging from onset information to texture and form. Evaluations of more than fifteen state-of-the-art models, in both zero-shot and fine-tuned settings, reveal pronounced modality gaps, unstable level-wise performance, and challenges in maintaining multilevel correctness. Fine-tuning substantially improves results across modalities while preserving general knowledge, positioning MSU-Bench as a robust foundation for future research in multimodal reasoning. The benchmark and code are available at https://github.com/Congren-Dai/MSU-Bench.

📄 PDF Abstract BibTeX arXiv:2511.20697

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningGeneral Knowledge

Similar Papers 제목 키워드 기반

SongEval: A Benchmark Dataset for Song Aesthetics Evaluation

2025-05-16 · Jixun Yao, Guobin Ma, Huixin Xue, Huakang Chen 외

Aesthetics serve as an implicit and important criterion in song generation tasks that reflect human perception beyond objective metrics. However, evaluating the aesthetics of generated songs remains a fundamental challen…

Evaluating Non-aligned Musical Score Transcriptions with MV2H

2019-06-03 · Andrew McLeod

The original MV2H metric was designed to evaluate systems which transcribe from an input audio (or MIDI) piece to a complete musical score. However, it requires both the transcribed score and the ground truth score to be…

Siamese Residual Neural Network for Musical Shape Evaluation in Piano Performance Assessment

2024-01-04 · Xiaoquan Li, Stephan Weiss, Yijun Yan, Yinhe Li 외

Understanding and identifying musical shape plays an important role in music education and performance assessment. To simplify the otherwise time- and cost-intensive musical shape evaluation, in this paper we explore how…

Universal Music Representations? Evaluating Foundation Models on World Music Corpora

2025-06-20 · Charilaos Papaioannou, Emmanouil Benetos, Alexandros Potamianos

Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of…

BenchmarkingFew-Shot LearningInformation RetrievalMusic Information Retrieval

Perception-Inspired Graph Convolution for Music Understanding Tasks

2024-05-15 · Emmanouil Karystinaios, Francesco Foscarin, Gerhard Widmer

We propose a new graph convolutional block, called MusGConv, specifically designed for the efficient processing of musical score data and motivated by general perceptual principles. It focuses on two fundamental dimensio…

Graph ClassificationGraph LearningLink PredictionNode Classification+1