paper-with-me

Papers

Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements

2020-12-07 · Kun Su, Xiulong Liu, Eli Shlizerman

We propose a novel system that takes as an input body movements of a musician playing a musical instrument and generates music in an unsupervised setting. Learning to generate multi-instrumental music from videos without labeling the instruments is a challenging problem. To achieve the transformation, we built a pipeline named 'Multi-instrumentalistNet' (MI Net). At its base, the pipeline learns a discrete latent representation of various instruments music from log-spectrogram using a Vector Quantized Variational Autoencoder (VQ-VAE) with multi-band residual blocks. The pipeline is then trained along with an autoregressive prior conditioned on the musician's body keypoints movements encoded by a recurrent neural network. Joint training of the prior with the body movements encoder succeeds in the disentanglement of the music into latent features indicating the musical components and the instrumental features. The latent space results in distributions that are clustered into distinct instruments from which new music can be generated. Furthermore, the VQ-VAE architecture supports detailed music generation with additional conditioning. We show that a Midi can further condition the latent space such that the pipeline will generate the exact content of the music being played by the instrument in the video. We evaluate MI Net on two datasets containing videos of 13 instruments and obtain generated music of reasonable audio quality, easily associated with the corresponding instrument, and consistent with the music audio content.

📄 PDF Abstract BibTeX arXiv:2012.03478

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementMusic Generation

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…
Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

2025-07-07 · Fathinah Izzati, Xinyue Li, Gus Xia arxiv

We propose Expotion (Facial Expression and Motion Control for Multimodal Music Generation), a generative model leveraging multimodal visual controls - specifically, human facial expressions and upper-body motion - as wel…

parameter-efficient fine-tuningText-to-Music Generation

GCDance: Genre-Controlled 3D Full Body Dance Generation Driven By Music

2025-02-25 · Xinran Liu, Xu Dong, Diptesh Kanojia, Wenwu Wang 외

Generating high-quality full-body dance sequences from music is a challenging task as it requires strict adherence to genre-specific choreography. Moreover, the generated sequences must be both physically realistic and p…

Rhythm

PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation

2026-01-22 · Jaekwon Im, Natalia Polouliakh, Taketo Akama arxiv

Dance-to-music generation aims to generate music that is aligned with dance movements. Existing approaches typically rely on body motion features extracted from a single human dancer and limited dance-to-music datasets, …

Music Generation

FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance Generation

2022-12-07 · ICCV 2023 1 · Ronghui Li, Junfan Zhao, Yachao Zhang, Mingyang Su 외

Generating full-body and multi-genre dance sequences from given music is a challenging task, due to the limitations of existing datasets and the inherent complexity of the fine-grained hand motion and dance genres. To ad…

Motion SynthesisRetrieval

MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation

2025-08-23 · Prerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta 외 arxiv

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion…