paper-with-me

홈 › Papers

Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal Space

2020-08-22 · Sicheng Zhao, Yaxian Li, Xingxu Yao, Wei-Zhi Nie, Pengfei Xu, Jufeng Yang, Kurt Keutzer

Both images and music can convey rich semantics and are widely used to induce specific emotions. Matching images and music with similar emotions might help to make emotion perceptions more vivid and stronger. Existing emotion-based image and music matching methods either employ limited categorical emotion states which cannot well reflect the complexity and subtlety of emotions, or train the matching model using an impractical multi-stage pipeline. In this paper, we study end-to-end matching between image and music based on emotions in the continuous valence-arousal (VA) space. First, we construct a large-scale dataset, termed Image-Music-Emotion-Matching-Net (IMEMNet), with over 140K image-music pairs. Second, we propose cross-modal deep continuous metric learning (CDCML) to learn a shared latent embedding space which preserves the cross-modal similarity relationship in the continuous matching space. Finally, we refine the embedding space by further preserving the single-modal emotion relationship in the VA spaces of both images and music. The metric learning in the embedding space and task regression in the label space are jointly optimized for both cross-modal matching and single-modal VA prediction. The extensive experiments conducted on IMEMNet demonstrate the superiority of CDCML for emotion-based image and music matching as compared to the state-of-the-art approaches.

📄 PDF Abstract BibTeX arXiv:2009.05103

Code (1)

linkAmy/IMEMNet 공식 구현

Tasks

Metric Learning

Similar Papers 제목 키워드 기반

MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions

2025-01-02 · Suhwan Choi, Kyu Won Kim, Myungjoo Kang

We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand t…

Emotion Embedding Spaces for Matching Music to Stories

2021-11-26 · Minz Won, Justin Salamon, Nicholas J. Bryan, Gautham J. Mysore 외

Content creators often use music to enhance their stories, as it can be a powerful tool to convey emotion. In this paper, our goal is to help creators find music to match the emotion of their story. We focus on text-base…

Cross-Modal RetrievalMetric LearningRetrieval

Exploring the Emotional Landscape of Music: An Analysis of Valence Trends and Genre Variations in Spotify Music Data

2023-10-29 · Shruti Dutta, Shashwat Mookherjee

This paper conducts an intricate analysis of musical emotions and trends using Spotify music data, encompassing audio features and valence scores extracted through the Spotipi API. Employing regression modeling, temporal…

regression

Emotion-Guided Image to Music Generation

2024-10-29 · Souraja Kundu, Saket Singh, Yuji Iwahori

Generating music from images can enhance various applications, including background music for photo slideshows, social media experiences, and video creation. This paper presents an emotion-guided image-to-music generatio…

Contrastive LearningMusic Generation

Do Picardy thirds smile? Tonal hierarchy and tonal valence: explicit and implicit measures

2021-03-17 · Neta B. Maimon, Dominique Lamy, Zohar Eitan

Western tonality provides a hierarchy of stability among melodic scale-degrees, from the maximally stable tonic to unstable chromatic notes. Tonal stability has been linked to emotion, yet systematic investigations of th…