paper-with-me

Papers

Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimodal Emotion Recognition

2023-11-18 · Dongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu Okumura

Multimodal emotion recognition aims to recognize emotions for each utterance of multiple modalities, which has received increasing attention for its application in human-machine interaction. Current graph-based methods fail to simultaneously depict global contextual features and local diverse uni-modal features in a dialogue. Furthermore, with the number of graph layers increasing, they easily fall into over-smoothing. In this paper, we propose a method for joint modality fusion and graph contrastive learning for multimodal emotion recognition (Joyful), where multimodality fusion, contrastive learning, and emotion recognition are jointly optimized. Specifically, we first design a new multimodal fusion mechanism that can provide deep interaction and fusion between the global contextual and uni-modal specific features. Then, we introduce a graph contrastive learning framework with inter-view and intra-view contrastive losses to learn more distinguishable representations for samples with different sentiments. Extensive experiments on three benchmark datasets indicate that Joyful achieved state-of-the-art (SOTA) performance compared to all baselines.

📄 PDF Abstract BibTeX arXiv:2311.11009

Code (1)

wykstc/MERC-main 공식 구현 pytorch

Tasks

Contrastive LearningEmotion RecognitionEmotion Recognition in ConversationFace SwappingMultimodal Emotion Recognition

Methods 이 논문이 사용한 방법론

Graph Contrastive Coding Graph Contrastive Coding is a self-supervised graph neural network pre-training framework to capture the universal network topological properties across multiple networks.…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

A Modality-Tailored Graph Modeling Framework for Urban Region Representation via Contrastive Learning

2025-09-28 · Yaya Zhao, Kaiqi Zhao, Zixuan Tang, Zhiyuan Liu 외 arxiv

Graph-based models have emerged as a powerful paradigm for modeling multimodal urban data and learning region representations for various downstream tasks. However, existing approaches face two major limitations. (1) The…

Contrastive LearningGraph Neural Network

The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment

2025-11-26 · Stefanos Koutoupis, Michaela Areti Zervou, Konstantinos Kontras, Maarten De Vos 외 arxiv

Learning joint representations across multiple modalities remains a central challenge in multimodal machine learning. Prevailing approaches predominantly operate in pairwise settings, aligning two modalities at a time. W…

Joint-Centric Dual Contrastive Alignment with Structure-Preserving and Information-Balanced Regularization

2026-04-17 · Habibeh Naderi, Behrouz Haji Soleimani, Stan Matwin arxiv

We propose HILBERT (HIerarchical Long-sequence Balanced Embedding with Reciprocal contrastive Training), a cross-attentive multimodal framework for learning document-level audio-text representations from long, segmented …

Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal Fusion

2019-11-18 · Sijie Mai, Haifeng Hu, Songlong Xing

Learning joint embedding space for various modalities is of vital importance for multimodal fusion. Mainstream modality fusion approaches fail to achieve this goal, leaving a modality gap which heavily affects cross-moda…

DecoderGraph Neural NetworkRepresentation LearningTranslation

Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

2025-10-13 · KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo 외 arxiv

Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Dif…

Audio captioning