paper-with-me

Papers

The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS

2025-10-21 · Brandon James Carone, Iran R. Roman, Pablo Ripollés arxiv

Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.

📄 PDF Abstract BibTeX arXiv:2510.19055

Code (0)

등록된 구현이 없습니다.

Tasks

Relational Reasoning

Similar Papers 제목 키워드 기반

Musical Training, but not Mere Exposure to Music, Drives the Emergence of Chroma Equivalence in Artificial Neural Networks

2026-02-20 · Lukas Grasse, Matthew S. Tata arxiv

Pitch is a fundamental aspect of auditory perception. Pitch perception is commonly described across two perceptual dimensions: pitch height is the sense that tones with varying frequencies seem to be higher or lower, and…

Self-Supervised LearningMusic Transcription

Attention but not musical training affects auditory streaming

2017-08-11

While musicians generally perform better than non-musicians in various auditory discrimination tasks, effects of specific instrumental training have received little attention. The effects of instrument-specific musical t…

Rhythm

Source Separation & Automatic Transcription for Music

2024-12-09 · Bradford Derby, Lucas Dunker, Samarth Galchar, Shashank Jarmale 외

Source separation is the process of isolating individual sounds in an auditory mixture of multiple sounds [1], and has a variety of applications ranging from speech enhancement and lyric transcription [2] to digital audi…

Music TranscriptionSpeech Enhancement

Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

2022-02-19 · Phoebe Chua, Dimos Makris, Dorien Herremans, Gemma Roig 외

Although media content is increasingly produced, distributed, and consumed in multiple combinations of modalities, how individual modalities contribute to the perceived emotion of a media item remains poorly understood. …

DescriptiveEmotion RecognitionFeature ImportanceMultimodal Emotion Recognition+1

End-to-end Topographic Auditory Models Replicate Signatures of Human Auditory Cortex

2025-09-28 · Haider Al-Tahan, Mayukh Deb, Jenelle Feather, N. Apurva Ratan Murty arxiv

The human auditory cortex is topographically organized. Neurons with similar response properties are spatially clustered, forming smooth maps for acoustic features such as frequency in early auditory areas, and modular r…