paper-with-me

홈 › Papers

HumMusQA: A Human-written Music Understanding QA Benchmark Dataset

2026-03-29 · Benno Weck, Pablo Puentes, Andrea Poltronieri, Satyajeet Prabhu, Dmitry Bogdanov arxiv

The evaluation of music understanding in Large Audio-Language Models (LALMs) requires a rigorously defined benchmark that truly tests whether models can perceive and interpret music, a standard that current data methodologies frequently fail to meet. This paper introduces a meticulously structured approach to music evaluation, proposing a new dataset of 320 hand-written questions curated and validated by experts with musical training, arguing that such focused, manual curation is superior for probing complex audio comprehension. To demonstrate the use of the dataset, we benchmark six state-of-the-art LALMs and additionally test their robustness to uni-modal shortcuts.

📄 PDF Abstract BibTeX arXiv:2603.27877

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

2023-11-16 · Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won 외

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural l…

Music CaptioningMusic GenerationRetrievalText to Audio Retrieval+1

Music Flamingo: Scaling Music Understanding in Audio Language Models

2025-11-13 · Sreyan Ghosh, Arushi Goel, Lasha Koroshinadze, Sang-gil Lee 외 arxiv

We introduce Music Flamingo, a novel large audio-language model designed to advance music (including song) understanding in foundational audio models. While audio-language research has progressed rapidly, music remains c…

Reinforcement Learning

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

2024-08-02 · Benno Weck, Ilaria Manco, Emmanouil Benetos, Elio Quinton 외

Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about…

Multimodal ReasoningMultiple-choiceMusic Question Answering

MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music

2024-02-15 · ZiHao Wang, Shuyu Li, Tao Zhang, Qi Wang 외

The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between …

Information RetrievalMusic Information Retrieval

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

2026-05-01 · Zuyao You, Zhesong Yu, Mingyu Liu, Bilei Zhu 외 arxiv

In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamlined encoder-decoder design of LLaVA, ena…

Reinforcement Learning