paper-with-me

홈 › Papers

Do Music Generation Models Encode Music Theory?

2024-10-01 · Megan Wei, Michael Freeman, Chris Donahue, Chen Sun

Music foundation models possess impressive music generation capabilities. When people compose music, they may infuse their understanding of music into their work, by using notes and intervals to craft melodies, chords to build progressions, and tempo to create a rhythmic feel. To what extent is this true of music generation models? More specifically, are fundamental Western music theory concepts observable within the "inner workings" of these models? Recent work proposed leveraging latent audio representations from music generation models towards music information retrieval tasks (e.g. genre classification, emotion recognition), which suggests that high-level musical characteristics are encoded within these models. However, probing individual music theory concepts (e.g. tempo, pitch class, chord quality) remains under-explored. Thus, we introduce SynTheory, a synthetic MIDI and audio music theory dataset, consisting of tempos, time signatures, notes, intervals, scales, chords, and chord progressions concepts. We then propose a framework to probe for these music theory concepts in music foundation models (Jukebox and MusicGen) and assess how strongly they encode these concepts within their internal representations. Our findings suggest that music theory concepts are discernible within foundation models and that the degree to which they are detectable varies by model size and layer.

📄 PDF Abstract BibTeX arXiv:2410.00872

Code (1)

brown-palm/syntheory 공식 구현 pytorch

Tasks

Emotion RecognitionGenre classificationInformation RetrievalMusic GenerationMusic Information Retrieval

Similar Papers 제목 키워드 기반

Composing Music with Grammar Argumented Neural Networks and Note-Level Encoding

2016-11-16 · Zheng Sun, Jiaqi Liu, Zewang Zhang, Jingwen Chen 외

Creating aesthetically pleasing pieces of art, including music, has been a long-term goal for artificial intelligence research. Despite recent successes of long-short term memory (LSTM) recurrent neural networks (RNNs) i…

Music Generation

MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding

2025-10-18 · Jingyue Huang, Zachary Novack, Phillip Long, Yupeng Hou 외 arxiv

Discrete representation learning has shown promising results across various domains, including generation and understanding in image, speech and language. Inspired by these advances, we propose MuseTok, a tokenization me…

Representation LearningEmotion RecognitionMusic Generation

JamBot: Music Theory Aware Chord Based Generation of Polyphonic Music with LSTMs

2017-11-21 · Gino Brunner, Yuyi Wang, Roger Wattenhofer, Jonas Wiesendanger

We propose a novel approach for the generation of polyphonic music based on LSTMs. We generate music in two steps. First, a chord LSTM predicts a chord progression based on a chord embedding. A second LSTM then generates…

MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core

2025-11-21 · Callie C. Liao, Duoduo Liao, Ellie L. Zhang arxiv

Recent advances in generative AI have made music generation a prominent research focus. However, many neural-based models rely on large datasets, raising concerns about copyright infringement and high-performance costs. …

Music Generation

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

2026-07-22 · Tieyao Zhang, Yuke Liu, Jiaxing Yu, Xinda Wu 외 arxiv

Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmen…

Music Generation