paper-with-me

Papers

JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata

2025-02-11 · Abhinaba Roy, Renhang Liu, Tongyu Lu, Dorien Herremans

We introduce JamendoMaxCaps, a large-scale music-caption dataset featuring over 362,000 freely licensed instrumental tracks from the renowned Jamendo platform. The dataset includes captions generated by a state-of-the-art captioning model, enhanced with imputed metadata. We also introduce a retrieval system that leverages both musical features and metadata to identify similar songs, which are then used to fill in missing metadata using a local large language model (LLLM). This approach allows us to provide a more comprehensive and informative dataset for researchers working on music-language understanding tasks. We validate this approach quantitatively with five different measurements. By making the JamendoMaxCaps dataset publicly available, we provide a high-quality resource to advance research in music-language understanding tasks such as music retrieval, multimodal representation learning, and generative music models.

📄 PDF Abstract BibTeX arXiv:2502.07461

Code (1)

amaai-lab/jamendomaxcaps 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelRepresentation LearningRetrieval

Similar Papers 제목 키워드 기반

LP-MusicCaps: LLM-Based Pseudo Music Captioning

2023-07-31 · Seungheon Doh, Keunwoo Choi, Jongpil Lee, Juhan Nam

Automatic music captioning, which generates natural language descriptions for given music tracks, holds significant potential for enhancing the understanding and organization of large volumes of musical data. Despite its…

Language ModelingLanguage ModellingLarge Language ModelMusic Captioning+2

MidiCaps: A large-scale MIDI dataset with text captions

2024-06-04 · Jan Melechovsky, Abhinaba Roy, Dorien Herremans

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combine…

Information RetrievalMusic Information Retrieval

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

2023-08-22 · Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, Ying Shan

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU…

Caption GenerationLarge Language ModelMultimodal Music GenerationMusic Captioning+2

Can Impressions of Music be Extracted from Thumbnail Images?

2025-01-05 · Takashi Harada, Takehiro Motomitsu, Katsuhiko Hayashi, Yusuke Sakai 외

In recent years, there has been a notable increase in research on machine learning models for music retrieval and generation systems that are capable of taking natural language sentences as inputs. However, there is a sc…

Retrieval

MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

2025-07-08 · Fathinah Izzati, Xinyue Li, Yuxuan Wu, Gus Xia

Humans can imagine various atmospheres and settings when listening to music, envisioning movie scenes that complement each piece. For example, slow, melancholic music might evoke scenes of heartbreak, while upbeat melodi…

Language ModelingLanguage ModellingMusic CaptioningMusic Generation