paper-with-me

홈 › Papers

MusicLIME: Explainable Multimodal Music Understanding

2024-09-16 · Theodoros Sotirou, Vassilis Lyberatos, Orfeas Menis Mastromichalakis, Giorgos Stamou

Multimodal models are critical for music understanding tasks, as they capture the complex interplay between audio and lyrics. However, as these models become more prevalent, the need for explainability grows-understanding how these systems make decisions is vital for ensuring fairness, reducing bias, and fostering trust. In this paper, we introduce MusicLIME, a model-agnostic feature importance explanation method designed for multimodal music models. Unlike traditional unimodal methods, which analyze each modality separately without considering the interaction between them, often leading to incomplete or misleading explanations, MusicLIME reveals how audio and lyrical features interact and contribute to predictions, providing a holistic view of the model's decision-making. Additionally, we enhance local explanations by aggregating them into global explanations, giving users a broader perspective of model behavior. Through this work, we contribute to improving the interpretability of multimodal music models, empowering users to make informed choices, and fostering more equitable, fair, and transparent music understanding systems.

📄 PDF Abstract BibTeX arXiv:2409.10496

Code (1)

iamtheo2000/musiclime 공식 구현

Tasks

Decision MakingFairnessFeature Importance

Similar Papers 제목 키워드 기반

Knowledge-based Multimodal Music Similarity

2023-06-21 · Andrea Poltronieri

Music similarity is an essential aspect of music retrieval, recommendation systems, and music analysis. Moreover, similarity is of vital interest for music experts, as it allows studying analogies and influences among co…

Recommendation SystemsRetrieval

DeformTune: A Deformable XAI Music Prototype for Non-Musicians

2025-07-31 · Ziqing Xu, Nick Bryan-Kinns arxiv

Many existing AI music generation tools rely on text prompts, complex interfaces, or instrument-like controls, which may require musical or technical knowledge that non-musicians do not possess. This paper introduces Def…

Music Generation

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

2024-08-02 · Benno Weck, Ilaria Manco, Emmanouil Benetos, Elio Quinton 외

Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about…

Multimodal ReasoningMultiple-choiceMusic Question Answering

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

2026-05-01 · Zuyao You, Zhesong Yu, Mingyu Liu, Bilei Zhu 외 arxiv

In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamlined encoder-decoder design of LLaVA, ena…

Reinforcement Learning

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

2025-02-18 · Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki 외

Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements. These improvements pri…