The Interpretation Gap in Text-to-Music Generation Models
Large-scale text-to-music generation models have significantly enhanced music creation capabilities, offering unprecedented creative freedom. However, their ability to collaborate effectively with human musicians remains limited. In this paper, we propose a framework to describe the musical interaction process, which includes expression, interpretation, and execution of controls. Following this framework, we argue that the primary gap between existing text-to-music models and musicians lies in the interpretation stage, where models lack the ability to interpret controls from musicians. We also propose two strategies to address this gap and call on the music information retrieval community to tackle the interpretation challenge to improve human-AI musical collaboration.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMusic GenerationMusic Information RetrievalRetrievalText-to-Music GenerationSimilar Papers 제목 키워드 기반
DanceChat: Large Language Model-Guided Music-to-Dance Generation
Music-to-dance generation aims to synthesize human dance motion conditioned on musical input. Despite recent progress, significant challenges remain due to the semantic gap between music and dance motion, as music offers…
Language ModelingLanguage ModellingLarge Language ModelMotion Synthesis+1Inspecting and Interacting with Meaningful Music Representations using VAE
Variational Autoencoders(VAEs) have already achieved great results on image generation and recently made promising progress on music generation. However, the generation process is still quite difficult to control in the …
DisentanglementImage GenerationMusic GenerationRhythmLearning long-term music representations via hierarchical contextual constraints
Learning symbolic music representations, especially disentangled representations with probabilistic interpretations, has been shown to benefit both music understanding and generation. However, most models are only applic…
Contrastive LearningDisentanglementRendering Music Performance With Interpretation Variations Using Conditional Variational RNN
Capturing and generating a wide variety of musical expression is important in music performance rendering, but current methods fail to model such a variation. This paper presents a music performance rendering method that…
DecoderMusic Performance RenderingJEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models
Music generation has attracted growing interest with the advancement of deep generative models. However, generating music conditioned on textual descriptions, known as text-to-music, remains challenging due to the comple…
Computational EfficiencyIn-Context LearningMusic GenerationText-to-Music Generation