paper-with-me

Papers

Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval

2024-07-28 · Zeyu Chen, Pengfei Zhang, Kai Ye, Wei Dong, Xin Feng, Yana Zhang

The burgeoning short video industry has accelerated the advancement of video-music retrieval technology, assisting content creators in selecting appropriate music for their videos. In self-supervised training for video-to-music retrieval, the video and music samples in the dataset are separated from the same video work, so they are all one-to-one matches. This does not match the real situation. In reality, a video can use different music as background music, and a music can be used as background music for different videos. Many videos and music that are not in a pair may be compatible, leading to false negative noise in the dataset. A novel inter-intra modal (II) loss is proposed as a solution. By reducing the variation of feature distribution within the two modalities before and after the encoder, II loss can reduce the model's overfitting to such noise without removing it in a costly and laborious way. The video-music retrieval framework, II-CLVM (Contrastive Learning for Video-Music Retrieval), incorporating the II Loss, achieves state-of-the-art performance on the YouTube8M dataset. The framework II-CLVTM shows better performance when retrieving music using multi-modal video information (such as text in videos). Experiments are designed to show that II loss can effectively alleviate the problem of false negative noise in retrieval tasks. Experiments also show that II loss improves various self-supervised and supervised uni-modal and cross-modal retrieval tasks, and can obtain good retrieval models with a small amount of training samples.

📄 PDF Abstract BibTeX arXiv:2407.19415

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningCross-Modal RetrievalRetrieval

Similar Papers 제목 키워드 기반

Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

2017-04-22 · Sungeun Hong, Woobin Im, Hyun S. Yang

Up to now, only limited research has been conducted on cross-modal retrieval of suitable music for a specified video or vice versa. Moreover, much of the existing research relies on metadata such as keywords, tags, or as…

Cross-Modal RetrievalRetrieval

Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs

2021-04-21 · Tingtian Li, Zixun Sun, Haoruo Zhang, Jin Li 외

Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval meth…

Pseudo LabelRetrievalTriplet

Audio-Visual Embedding for Cross-Modal MusicVideo Retrieval through Supervised Deep CCA

2019-08-10 · Donghuo Zeng, Yi Yu, Keizo Oyama

Deep learning has successfully shown excellent performance in learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal st…

audio-visual learningRetrievalVideo Retrieval

Generative Disco: Text-to-Video Generation for Music Visualization

2023-04-17 · Vivian Liu, Tao Long, Nathan Raw, Lydia Chilton

Visuals can enhance our experience of music, owing to the way they can amplify the emotions and messages conveyed within it. However, creating music visualization is a complex, time-consuming, and resource-intensive proc…

Text-to-Video GenerationVideo Generation

Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

2024-12-12 · Baisen Wang, Le Zhuo, Zhaokai Wang, Chenxi Bao 외

Multimodal music generation aims to produce music from diverse input modalities, including text, videos, and images. Existing methods use a common embedding space for multimodal fusion. Despite their effectiveness in oth…

cross-modal alignmentMultimodal Music GenerationMusic GenerationRetrieval