paper-with-me

Papers

Learning Music Sequence Representation from Text Supervision

2023-05-31 · Tianyu Chen, Yuan Xie, Shuai Zhang, Shaohan Huang, Haoyi Zhou, JianXin Li

Music representation learning is notoriously difficult for its complex human-related concepts contained in the sequence of numerical signals. To excavate better MUsic SEquence Representation from labeled audio, we propose a novel text-supervision pre-training method, namely MUSER. MUSER adopts an audio-spectrum-text tri-modal contrastive learning framework, where the text input could be any form of meta-data with the help of text templates while the spectrum is derived from an audio sequence. Our experiments reveal that MUSER could be more flexibly adapted to downstream tasks compared with the current data-hungry pre-training method, and it only requires 0.056% of pre-training data to achieve the state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2305.19602

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningRepresentation Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Vector Quantized Contrastive Predictive Coding for Template-based Music Generation

2020-04-21 · Gaëtan Hadjeres, Léopold Crestel

In this work, we propose a flexible method for generating variations of discrete sequences in which tokens can be grouped into basic units, like sentences in a text or bars in music. More precisely, given a template sequ…

Music Generation

Learning music audio representations via weak language supervision

2021-12-08 · Ilaria Manco, Emmanouil Benetos, Elio Quinton, Gyorgy Fazekas

Audio representations for music information retrieval are typically learned via supervised learning in a task-specific fashion. Although effective at producing state-of-the-art results, this scheme lacks flexibility with…

Audio ClassificationInformation RetrievalMusic Information RetrievalRetrieval

Video-to-Music Recommendation using Temporal Alignment of Segments

2023-06-12 · Laure Prétet, Gaël Richard, Clément Souchier, Geoffroy Peeters

We study cross-modal recommendation of music tracks to be used as soundtracks for videos. This problem is known as the music supervision task. We build on a self-supervised system that learns a content association betwee…

Music Recommendation

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

2026-08-10 · Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji 외 arxiv

Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise…

Text-to-Music GenerationAudio Generation

Semi-Supervised Contrastive Learning of Musical Representations

2024-07-18 · Julien Guinot, Elio Quinton, György Fazekas

Despite the success of contrastive learning in Music Information Retrieval, the inherent ambiguity of contrastive self-supervision presents a challenge. Relying solely on augmentation chains and self-supervised positive …

Contrastive LearningInformation RetrievalMusic Information RetrievalTransfer Learning