paper-with-me

홈 › Papers

Discovering Interpretable Concepts in Large Generative Music Models

2025-05-18 · Nikhil Singh, Manuel Cherep, Pattie Maes

The fidelity with which neural networks can now generate content such as music presents a scientific opportunity: these systems appear to have learned implicit theories of the structure of such content through statistical learning alone. This could offer a novel lens on theories of human-generated media. Where these representations align with traditional constructs (e.g. chord progressions in music), they demonstrate how these can be inferred from statistical regularities. Where they diverge, they highlight potential limits in our theoretical frameworks -- patterns that we may have overlooked but that nonetheless hold significant explanatory power. In this paper, we focus on the specific case of music generators. We introduce a method to discover musical concepts using sparse autoencoders (SAEs), extracting interpretable features from the residual stream activations of a transformer model. We evaluate this approach by extracting a large set of features and producing an automatic labeling and evaluation pipeline for them. Our results reveal both familiar musical concepts and counterintuitive patterns that lack clear counterparts in existing theories or natural language altogether. Beyond improving model transparency, our work provides a new empirical tool that might help discover organizing principles in ways that have eluded traditional methods of analysis and synthesis.

📄 PDF Abstract BibTeX arXiv:2505.18186

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ProGress: Structured Music Generation via Graph Diffusion and Hierarchical Music Analysis

2025-10-11 · Stephen Ni-Hahn, Chao Péter Yang, Mingchen Ma, Cynthia Rudin 외 arxiv

Artificial Intelligence (AI) for music generation is undergoing rapid developments, with recent symbolic models leveraging sophisticated deep learning and diffusion model algorithms. One drawback with existing models is …

Music Generation

Learning Interpretable Musical Compositional Rules and Traces

2016-06-17 · Haizi Yu, Lav R. Varshney, Guy E. Garnett, Ranjitha Kumar

Throughout music history, theorists have identified and documented interpretable rules that capture the decisions of composers. This paper asks, "Can a machine behave like a music theorist?" It presents MUS-ROVER, a self…

Self-Learning

Semantic-Aware Interpretable Multimodal Music Auto-Tagging

2025-05-22 · Andreas Patakis, Vassilis Lyberatos, Spyridon Kantarelis, Edmund Dervakos 외

Music auto-tagging is essential for organizing and discovering music in extensive digital libraries. While foundation models achieve exceptional performance in this domain, their outputs often lack interpretability, limi…

Decision MakingMusic Auto-TaggingMusic Tagging

Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders

2025-10-27 · Nathan Paek, Yongyi Zang, Qihui Yang, Randal Leistikow arxiv

While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature requires compression that obscures semant…

Audio GenerationMusic Generation

JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters Tuning

2024-06-18 · BoYu Chen, Peike Li, Yao Yao, Alex Wang

Large models for text-to-music generation have achieved significant progress, facilitating the creation of high-quality and varied musical compositions from provided text prompts. However, input text prompts may not prec…

Music GenerationText-to-Music Generation