paper-with-me

Papers

MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions

2025-01-02 · Suhwan Choi, Kyu Won Kim, Myungjoo Kang

We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand the Image-Music-Emotion-Matching-Net (IMEMNet) dataset, creating IMEMNet-C which includes 24,756 images and 25,944 music clips with corresponding musical captions. We employ multimodal matching scores based on the continuous valence (emotional positivity) and arousal (emotional intensity) values. This continuous matching score allows for random sampling of image-music pairs during training by computing similarity scores from the valence-arousal values across different modalities. Consequently, the proposed approach achieves state-of-the-art performance in valence-arousal prediction tasks. Furthermore, the framework demonstrates its efficacy in various zeroshot tasks, highlighting the potential of valence and arousal predictions in downstream applications.

📄 PDF Abstract BibTeX arXiv:2501.01094

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal Emotion Recognition for One-Minute-Gradual Emotion Challenge

2018-05-03 · Ziqi Zheng, Chenjie Cao, Xingwei Chen, Guoqiang Xu

The continuous dimensional emotion modelled by arousal and valence can depict complex changes of emotions. In this paper, we present our works on arousal and valence predictions for One-Minute-Gradual (OMG) Emotion Chall…

Emotion RecognitionMultimodal Emotion Recognition

Using large language models to estimate features of multi-word expressions: Concreteness, valence, arousal

2024-08-16 · Gonzalo Martínez, Juan Diego Molero, Sandra González, Javier Conde 외

This study investigates the potential of large language models (LLMs) to provide accurate estimates of concreteness, valence and arousal for multi-word expressions. Unlike previous artificial intelligence (AI) methods, L…

Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach

2026-03-13 · Elena Ryumina, Maxim Markitantov, Alexandr Axyonov, Dmitry Ryumin 외 arxiv

Continuous emotion recognition in terms of valence and arousal under in-the-wild (ITW) conditions remains a challenging problem due to large variations in appearance, head pose, illumination, occlusions, and subject-spec…

Emotion Recognition

Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction

2025-05-08 · Von Ralph Dane Marquez Herbuela, Yukie Nagai

Human emotional expression emerges through coordinated vocal, facial, and gestural signals. While speech face alignment is well established, the broader dynamics linking emotionally expressive speech to regional facial a…

Face Alignment

Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition

2024-03-20 · R. Gnana Praveen, Jahangir Alam

Though multimodal emotion recognition has achieved significant progress over recent years, the potential of rich synergic relationships across the modalities is not fully exploited. In this paper, we introduce Recursive …

Emotion RecognitionMultimodal Emotion Recognition