Performance Conditioning for Diffusion-Based Multi-Instrument Music Synthesis
Generating multi-instrument music from symbolic music representations is an important task in Music Information Retrieval (MIR). A central but still largely unsolved problem in this context is musically and acoustically informed control in the generation process. As the main contribution of this work, we propose enhancing control of multi-instrument synthesis by conditioning a generative model on a specific performance and recording environment, thus allowing for better guidance of timbre and style. Building on state-of-the-art diffusion-based music generative models, we introduce performance conditioning - a simple tool indicating the generative model to synthesize music with style and timbre of specific instruments taken from specific performances. Our prototype is evaluated using uncurated performances with diverse instrumentation and achieves state-of-the-art FAD realism scores while allowing novel timbre and style control. Our project page, including samples and demonstrations, is available at benadar293.github.io/midipm
Code (0)
등록된 구현이 없습니다.
Tasks
FADInformation RetrievalMusic Information RetrievalRetrievalSimilar Papers 제목 키워드 기반
MuScriptor: An Open Model for Multi-Instrument Music Transcription
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models gen…
Multi-instrument Music TranscriptionReinforcement LearningTimbre transfer using image-to-image denoising diffusion implicit models
Timbre transfer techniques aim at converting the sound of a musical piece generated by one instrument into the same one as if it was played by another instrument, while maintaining as much as possible the content in term…
DenoisingImage DenoisingRenderBox: Expressive Performance Rendering with Text Control
Expressive music performance rendering involves interpreting symbolic scores with variations in timing, dynamics, articulation, and instrument-specific techniques, resulting in performances that capture musical can emoti…
DiversityFADMusic Performance RenderingMulti-Instrumentalist Net: Unsupervised Generation of Music from Body Movements
We propose a novel system that takes as an input body movements of a musician playing a musical instrument and generates music in an unsupervised setting. Learning to generate multi-instrumental music from videos without…
DisentanglementMusic GenerationGTR-CTRL: Instrument and Genre Conditioning for Guitar-Focused Music Generation with Transformers
Recently, symbolic music generation with deep learning techniques has witnessed steady improvements. Most works on this topic focus on MIDI representations, but less attention has been paid to symbolic music generation u…
Genre classificationMusic GenerationRhythm