Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal Space
Both images and music can convey rich semantics and are widely used to induce specific emotions. Matching images and music with similar emotions might help to make emotion perceptions more vivid and stronger. Existing emotion-based image and music matching methods either employ limited categorical emotion states which cannot well reflect the complexity and subtlety of emotions, or train the matching model using an impractical multi-stage pipeline. In this paper, we study end-to-end matching between image and music based on emotions in the continuous valence-arousal (VA) space. First, we construct a large-scale dataset, termed Image-Music-Emotion-Matching-Net (IMEMNet), with over 140K image-music pairs. Second, we propose cross-modal deep continuous metric learning (CDCML) to learn a shared latent embedding space which preserves the cross-modal similarity relationship in the continuous matching space. Finally, we refine the embedding space by further preserving the single-modal emotion relationship in the VA spaces of both images and music. The metric learning in the embedding space and task regression in the label space are jointly optimized for both cross-modal matching and single-modal VA prediction. The extensive experiments conducted on IMEMNet demonstrate the superiority of CDCML for emotion-based image and music matching as compared to the state-of-the-art approaches.
Code (1)
Tasks
Metric LearningSimilar Papers 제목 키워드 기반
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand t…
Emotion Embedding Spaces for Matching Music to Stories
Content creators often use music to enhance their stories, as it can be a powerful tool to convey emotion. In this paper, our goal is to help creators find music to match the emotion of their story. We focus on text-base…
Cross-Modal RetrievalMetric LearningRetrievalExploring the Emotional Landscape of Music: An Analysis of Valence Trends and Genre Variations in Spotify Music Data
This paper conducts an intricate analysis of musical emotions and trends using Spotify music data, encompassing audio features and valence scores extracted through the Spotipi API. Employing regression modeling, temporal…
regressionEmotion-Guided Image to Music Generation
Generating music from images can enhance various applications, including background music for photo slideshows, social media experiences, and video creation. This paper presents an emotion-guided image-to-music generatio…
Contrastive LearningMusic GenerationDo Picardy thirds smile? Tonal hierarchy and tonal valence: explicit and implicit measures
Western tonality provides a hierarchy of stability among melodic scale-degrees, from the maximally stable tonic to unstable chromatic notes. Tonal stability has been linked to emotion, yet systematic investigations of th…