paper-with-me

홈 › Papers

Emotion Embedding Spaces for Matching Music to Stories

2021-11-26 · Minz Won, Justin Salamon, Nicholas J. Bryan, Gautham J. Mysore, Xavier Serra

Content creators often use music to enhance their stories, as it can be a powerful tool to convey emotion. In this paper, our goal is to help creators find music to match the emotion of their story. We focus on text-based stories that can be auralized (e.g., books), use multiple sentences as input queries, and automatically retrieve matching music. We formalize this task as a cross-modal text-to-music retrieval problem. Both the music and text domains have existing datasets with emotion labels, but mismatched emotion vocabularies prevent us from using mood or emotion annotations directly for matching. To address this challenge, we propose and investigate several emotion embedding spaces, both manually defined (e.g., valence/arousal) and data-driven (e.g., Word2Vec and metric learning) to bridge this gap. Our experiments show that by leveraging these embedding spaces, we are able to successfully bridge the gap between modalities to facilitate cross modal retrieval. We show that our method can leverage the well established valence-arousal space, but that it can also achieve our goal via data-driven embedding spaces. By leveraging data-driven embeddings, our approach has the potential of being generalized to other retrieval tasks that require broader or completely different vocabularies.

📄 PDF Abstract BibTeX arXiv:2111.13468

Code (1)

minzwon/text2music-emotion-embedding 공식 구현 pytorch

Tasks

Cross-Modal RetrievalMetric LearningRetrieval

Similar Papers 제목 키워드 기반

Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal Space

2020-08-22 · Sicheng Zhao, Yaxian Li, Xingxu Yao, Wei-Zhi Nie 외

Both images and music can convey rich semantics and are widely used to induce specific emotions. Matching images and music with similar emotions might help to make emotion perceptions more vivid and stronger. Existing em…

Metric Learning

MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions

2025-01-02 · Suhwan Choi, Kyu Won Kim, Myungjoo Kang

We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand t…

Comparison and Analysis of Deep Audio Embeddings for Music Emotion Recognition

2021-04-13 · Eunjeong Koh, Shlomo Dubnov

Emotion is a complicated notion present in music that is hard to capture even with fine-tuned feature engineering. In this paper, we investigate the utility of state-of-the-art pre-trained deep audio embedding methods to…

Emotion RecognitionFeature EngineeringMusic Emotion Recognition

SyMuPe: Affective and Controllable Symbolic Music Performance

2025-11-05 · Ilya Borovik, Dmitrii Gavrilev, Vladimir Viro arxiv

Emotions are fundamental to the creation and perception of music performances. However, achieving human-like expression and emotion through machine learning models for performance rendering remains a challenging task. In…

Modeling Tension in Stories via Commonsense Reasoning and Emotional Word Embeddings

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Dramatic tension is crucial for generating interesting stories. This paper aims to model dramatic tension from a story text using neural commonsense-reasoning language models and emotional word embeddings. We also propos…

Word Embeddings