paper-with-me

홈 › Papers

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

2026-05-01 · Zuyao You, Zhesong Yu, Mingyu Liu, Bilei Zhu, Yuan Wan, Zuxuan Wu arxiv

In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamlined encoder-decoder design of LLaVA, enabling effective cross-modal learning between music and language. By incorporating audio encoders in a mixture-of-experts manner, GaMMA effectively unifies both time-series and non-time-series music understanding tasks within one set of parameters. Our approach combines carefully curated datasets at scale with a progressive training pipeline, effectively pushing the boundaries of music understanding via pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL). To comprehensively assess both temporal and non-temporal capability of music LMMs, we introduce MusicBench, the largest music-oriented benchmark, comprising 3,739 human-curated multiple-choice questions covering diverse aspects of musical understanding. Extensive experiments demonstrate that GaMMA establishes new SoTA in the music domain, achieving 79.1% accuracy on MuchoMusic, 79.3% on MusicBench-Temporal, and 81.3% on MusicBench-Global, consistently outperforming previous methods.

📄 PDF Abstract BibTeX arXiv:2605.00371

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

GACA-DiT: Diffusion-based Dance-to-Music Generation with Genre-Adaptive Rhythm and Context-Aware Alignment

2025-10-28 · Jinting Wang, Chenxing Li, Li Liu arxiv

Dance-to-music (D2M) generation aims to automatically compose music that is rhythmically and temporally aligned with dance movements. Existing methods typically rely on coarse rhythm embeddings, such as global motion fea…

Music Generation

Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation

2024-07-29 · Junda Wu, Zachary Novack, Amit Namburi, Jiaheng Dai 외

Existing music captioning methods are limited to generating concise global descriptions of short music clips, which fail to capture fine-grained musical characteristics and time-aware musical changes. To address these li…

Music CaptioningMusic Generation

Modeling Temporal Lobe Epilepsy during Music Large-Scale Form Perception using the Impulse Pattern Formulation (IPF) Brain Mode

2023-10-05 · Rolf Bader

Musical large-scale form is investigated using an Electronic Dance Music (EDM) piece fed into a Finite-Difference Time Domain (FDTD) physical model of the cochlear which again inputs into an Impulse-Pattern Formulation (…

EEGFormRhythm

Rhythm is a Dancer: Music-Driven Motion Synthesis with Global Structure

2021-11-23 · Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman, Daniel Cohen-Or 외

Synthesizing human motion with a global structure, such as a choreography, is a challenging task. Existing methods tend to concentrate on local smooth pose transitions and neglect the global context or the theme of the m…

Motion SynthesisRhythm

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

2026-05-28 · Daeyong Kwon, Qiyu Wu, Shinobu Kuriya, Junghyun Koo 외 arxiv

Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temporal regions of the audio remains undere…