paper-with-me

홈 › Papers

Assessing Factual Music Comprehension in Large Audio Language Models

2025-11-02 · Daniel Chenyu Lin, Michael Freeman, John Thickstun arxiv

Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empirical evidence that assessment of LALMs using the popular MusicQA dataset fails to measure whether a model's responses about music are factually correct, and (2) develop a new protocol for assessing the music comprehension capabilities of LALMs. Specifically, we propose an evaluation protocol that prompts a LALM for factually verifiable information, and parses its open-ended response into a structured format that can be objectively assessed using Precision, Recall, and F1 scores. Using this protocol, we define a benchmark consisting of six factual information retrieval tasks defined on three diverse datasets: MusicNet, the Free Music Archive, and OverClocked ReMix. We benchmark nine recent LALMs, including frontier models like Gemini and the latest open models like Music Flamingo, and release the suite of evaluation scripts at https://github.com/DCL2004/LALM-Eval to facilitate benchmarking of new LALMs.

📄 PDF Abstract BibTeX arXiv:2511.05550

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language QueriesInformation Retrieval

Similar Papers 제목 키워드 기반

AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

2024-02-12 · Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu 외

Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-centric interaction capabilities has impeded…

2kAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarking+3

HumMusQA: A Human-written Music Understanding QA Benchmark Dataset

2026-03-29 · Benno Weck, Pablo Puentes, Andrea Poltronieri, Satyajeet Prabhu 외 arxiv

The evaluation of music understanding in Large Audio-Language Models (LALMs) requires a rigorously defined benchmark that truly tests whether models can perceive and interpret music, a standard that current data methodol…

Audio-FLAN: A Preliminary Release

2025-02-23 · Liumeng Xue, Ziya Zhou, Jiahao Pan, Zixuan Li 외

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tas…

Zero-Shot Learning

Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance

2024-09-23 · Yuanchao Li, Azalea Gui, Dimitra Emmanouilidou, Hannes Gamper

The complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a…

Emotion RecognitionFADMusic Emotion RecognitionMusic Generation

Introducing a Framework for the Evaluation of Music Detection Tools

2014-05-01 · LREC 2014 5 · Paula Lopez-Otero, Docio-Fern, Laura ez, Carmen Garcia-Mateo

The huge amount of multimedia information available nowadays makes its manual processing prohibitive, requiring tools for automatic labelling of these contents. This paper describes a framework for assessing a music dete…

Information RetrievalMusic Information RetrievalRetrieval