paper-with-me

홈 › Papers

HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization

2025-03-03 · CVPR 2025 1 · Zitang Zhou, Ke Mei, Yu Lu, Tianyi Wang, Fengyun Rao

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional alignment, thematic coherence, and cultural relevance. We propose a multi-step human-machine collaborative framework for efficient annotation, combining human insights with machine-generated descriptions to identify key transitions and assess alignment across multiple dimensions. Additionally, we introduce a novel evaluation framework with tasks and metrics to assess the multi-dimensional alignment of video and music, including rhythm, emotion, theme, and cultural context. Our extensive experiments demonstrate that HarmonySet, along with the proposed evaluation framework, significantly improves the ability of multimodal models to capture and analyze the intricate relationships between video and music.

📄 PDF Abstract BibTeX arXiv:2503.01725

Code (0)

등록된 구현이 없습니다.

Tasks

Rhythm

Similar Papers 제목 키워드 기반

MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

2025-07-08 · Fathinah Izzati, Xinyue Li, Yuxuan Wu, Gus Xia

Humans can imagine various atmospheres and settings when listening to music, envisioning movie scenes that complement each piece. For example, slow, melancholic music might evoke scenes of heartbreak, while upbeat melodi…

Language ModelingLanguage ModellingMusic CaptioningMusic Generation

A Comprehensive Survey on Generative AI for Video-to-Music Generation

2025-02-18 · Shulei Ji, Songruoyao Wu, ZiHao Wang, Shuyu Li 외

The burgeoning growth of video-to-music generation can be attributed to the ascendancy of multimodal generative models. However, there is a lack of literature that comprehensively combs through the work in this field. To…

Music Generation

Cross-Modal Learning for Music-to-Music-Video Description Generation

2025-03-14 · Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Zhi Zhong 외

Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models has opened a promising pathway for musi…

Video DescriptionVideo Generation

DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning

2025-02-18 · Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Hiromi Wakaki 외

Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements. These improvements pri…

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

2026-05-01 · Zuyao You, Zhesong Yu, Mingyu Liu, Bilei Zhu 외 arxiv

In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamlined encoder-decoder design of LLaVA, ena…

Reinforcement Learning