paper-with-me

Music Question Answering

1개 벤치마크 · 논문 7편 · 이 태스크의 논문 보기 →

Benchmarks

MusicQA

결과 3개

Most implemented

Listen, Think, and Understand

2023-05-18 · 구현 1개

Papers

Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding

2026-08-26 · Tianle Wang, Xinyi Tong, Liangke Zhao, Jishang Chen 외 arxiv

Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the \emph{Dissonance Spectrum} (DS), a nonn…

Music Question AnsweringEmotion Recognition

ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering

2025-12-05 · Daeyong Kwon, SeungHeon Doh, Juhan Nam arxiv

Recent advances in large language models (LLMs) have transformed open-domain question answering, yet their effectiveness in music-related reasoning remains limited due to sparse music knowledge in pretraining data. While…

Open-Domain Question AnsweringMusic Question AnsweringInformation Retrieval

MUST-RAG: MUSical Text Question Answering with Retrieval Augmented Generation

2025-07-31 · Daeyong Kwon, SeungHeon Doh, Juhan Nam arxiv

Recent advancements in Large language models (LLMs) have demonstrated remarkable capabilities across diverse domains. While they exhibit strong zero-shot performance on various tasks, LLMs' effectiveness in music-related…

Computational EfficiencyMusic Question AnsweringDomain Adaptation

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

2024-08-02 · Benno Weck, Ilaria Manco, Emmanouil Benetos, Elio Quinton 외

Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about…

Multimodal ReasoningMultiple-choiceMusic Question Answering

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

2023-08-22 · Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, Ying Shan

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU…

Caption GenerationLarge Language ModelMultimodal Music GenerationMusic Captioning+2

Listen, Think, and Understand

2023-05-18 · Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky 외

The ability of artificial intelligence (AI) systems to perceive and comprehend audio signals is crucial for many applications. Although significant progress has been made in this area since the development of AudioSet, m…

Language ModellingLarge Language ModelMusic Question Answering

전체 7편 보기 →