paper-with-me

홈 › Papers

The Interpretation Gap in Text-to-Music Generation Models

2024-07-14 · Yongyi Zang, Yixiao Zhang

Large-scale text-to-music generation models have significantly enhanced music creation capabilities, offering unprecedented creative freedom. However, their ability to collaborate effectively with human musicians remains limited. In this paper, we propose a framework to describe the musical interaction process, which includes expression, interpretation, and execution of controls. Following this framework, we argue that the primary gap between existing text-to-music models and musicians lies in the interpretation stage, where models lack the ability to interpret controls from musicians. We also propose two strategies to address this gap and call on the music information retrieval community to tackle the interpretation challenge to improve human-AI musical collaboration.

📄 PDF Abstract BibTeX arXiv:2407.10328

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMusic GenerationMusic Information RetrievalRetrievalText-to-Music Generation

Similar Papers 제목 키워드 기반

DanceChat: Large Language Model-Guided Music-to-Dance Generation

2025-06-12 · Qing Wang, Xiaohang Yang, Yilan Dong, Naveen Raj Govindaraj 외

Music-to-dance generation aims to synthesize human dance motion conditioned on musical input. Despite recent progress, significant challenges remain due to the semantic gap between music and dance motion, as music offers…

Language ModelingLanguage ModellingLarge Language ModelMotion Synthesis+1

Inspecting and Interacting with Meaningful Music Representations using VAE

2019-04-18 · Ruihan Yang, Tianyao Chen, Yiyi Zhang, Gus Xia

Variational Autoencoders(VAEs) have already achieved great results on image generation and recently made promising progress on music generation. However, the generation process is still quite difficult to control in the …

DisentanglementImage GenerationMusic GenerationRhythm

Learning long-term music representations via hierarchical contextual constraints

2022-02-13 · Shiqi Wei, Gus Xia

Learning symbolic music representations, especially disentangled representations with probabilistic interpretations, has been shown to benefit both music understanding and generation. However, most models are only applic…

Contrastive LearningDisentanglement

Rendering Music Performance With Interpretation Variations Using Conditional Variational RNN

2019-11-04 · ISMIR 2019 11 · Akira Maezawa, Kazuhiko Yamamoto, Takuya Fujishima

Capturing and generating a wide variety of musical expression is important in music performance rendering, but current methods fail to model such a variation. This paper presents a music performance rendering method that…

DecoderMusic Performance Rendering

JEN-1: Text-Guided Universal Music Generation with Omnidirectional Diffusion Models

2023-08-09 · Peike Li, BoYu Chen, Yao Yao, Yikai Wang 외

Music generation has attracted growing interest with the advancement of deep generative models. However, generating music conditioned on textual descriptions, known as text-to-music, remains challenging due to the comple…

Computational EfficiencyIn-Context LearningMusic GenerationText-to-Music Generation