paper-with-me

Papers Music Captioning

“Music Captioning” 태그가 달린 논문 14편 · 필터 해제

MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

2025-07-08 · Fathinah Izzati, Xinyue Li, Yuxuan Wu, Gus Xia

Humans can imagine various atmospheres and settings when listening to music, envisioning movie scenes that complement each piece. For example, slow, melancholic music might evoke scenes of heartbreak, while upbeat melodi…

Language ModelingLanguage ModellingMusic CaptioningMusic Generation

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning

2025-06-18 · Anuradha Chopra, Abhinaba Roy, Dorien Herremans

Detailed captions that accurately reflect the characteristics of a music piece can enrich music databases and drive forward research in music AI. This paper introduces a multi-task music captioning model, SonicVerse, tha…

Caption GenerationDescriptiveKey DetectionLarge Language Model+2

SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling

2025-06-17 · Tawsif Ahmed, Andrej Radonjic, Gollam Rabby

We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song. To the best of our knowledge, there are no open-source high-quality dataset representing popular and well-known songs for generative mu…

Music CaptioningMusic ModelingSinging Voice Synthesis

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

2025-06-14 · Yinghao Ma, Siyou Li, Juntao Yu, Emmanouil Benetos 외

Recent advances in audio-text large language models (LLMs) have opened new possibilities for music understanding and generation. However, existing benchmarks are limited in scope, often relying on simplified tasks or mul…

Beat TrackingGenre classificationInformation RetrievalInstruction Following+5

Do Captioning Metrics Reflect Music Semantic Alignment?

2024-11-18 · Jinwoo Lee, Kyogu Lee

Music captioning has emerged as a promising task, fueled by the advent of advanced language generation models. However, the evaluation of music captioning relies heavily on traditional metrics such as BLEU, METEOR, and R…

Music CaptioningText Generation

Evaluation of pretrained language models on music understanding

2024-09-17 · Yannis Vasilakis, Rachel Bittner, Johan Pauwels

Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the …

Music CaptioningNegationSensitivityText to Audio Retrieval+1

Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation

2024-07-29 · Junda Wu, Zachary Novack, Amit Namburi, Jiaheng Dai 외

Existing music captioning methods are limited to generating concise global descriptions of short music clips, which fail to capture fine-grained musical characteristics and time-aware musical changes. To address these li…

Music CaptioningMusic Generation

The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

2023-11-16 · Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won 외

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural l…

Music CaptioningMusic GenerationRetrievalText to Audio Retrieval+1

MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response

2023-09-15 · Zihao Deng, Yinghao Ma, Yudong Liu, Rongchen Guo 외

Large Language Models (LLMs) have shown immense potential in multimodal applications, yet the convergence of textual and musical domains remains not well-explored. To address this gap, we present MusiLingo, a novel syste…

Caption GenerationLanguage ModellingMusic Captioning

Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning

2023-08-22 · Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, Ying Shan

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU…

Caption GenerationLarge Language ModelMultimodal Music GenerationMusic Captioning+2

LP-MusicCaps: LLM-Based Pseudo Music Captioning

2023-07-31 · Seungheon Doh, Keunwoo Choi, Jongpil Lee, Juhan Nam

Automatic music captioning, which generates natural language descriptions for given music tracks, holds significant potential for enhancing the understanding and organization of large volumes of musical data. Despite its…

Language ModelingLanguage ModellingLarge Language ModelMusic Captioning+2

ALCAP: Alignment-Augmented Music Captioner

2022-12-21 · Zihao He, Weituo Hao, Wei-Tsung Lu, Changyou Chen 외

Music captioning has gained significant attention in the wake of the rising prominence of streaming media platforms. Traditional approaches often prioritize either the audio or lyrics aspect of the music, inadvertently i…

Contrastive LearningMusic CaptioningRetrieval

MusCaps: Generating Captions for Music Audio

2021-04-24 · Ilaria Manco, Emmanouil Benetos, Elio Quinton, Gyorgy Fazekas

Content-based music information retrieval has seen rapid progress with the adoption of deep learning. Current approaches to high-level music description typically make use of classification models, such as in auto-taggin…

Audio captioningClassificationCross-Modal RetrievalDecoder+3

Towards Music Captioning: Generating Music Playlist Descriptions

2016-08-17 · Keunwoo Choi, George Fazekas, Brian McFee, Kyunghyun Cho 외

Descriptions are often provided along with recommendations to help users' discovery. Recommending automatically generated music playlists (e.g. personalised playlists) introduces the problem of generating descriptions. I…

Music Captioning
1–14 / 14