paper-with-me

홈 › Papers

AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations

2024-04-12 · Sheng Wu, Jiaxing Liu, Longbiao Wang, Dongxiao He, Xiaobao Wang, Jianwu Dang

Emotion Recognition in Conversations (ERC) is a popular task in natural language processing, which aims to recognize the emotional state of the speaker in conversations. While current research primarily emphasizes contextual modeling, there exists a dearth of investigation into effective multimodal fusion methods. We propose a novel framework called AIMDiT to solve the problem of multimodal fusion of deep features. Specifically, we design a Modality Augmentation Network which performs rich representation learning through dimension transformation of different modalities and parameter-efficient inception block. On the other hand, the Modality Interaction Network performs interaction fusion of extracted inter-modal features and intra-modal features. Experiments conducted using our AIMDiT framework on the public benchmark dataset MELD reveal 2.34% and 2.87% improvements in terms of the Acc-7 and w-F1 metrics compared to the state-of-the-art (SOTA) models.

📄 PDF Abstract BibTeX arXiv:2407.00743

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionRepresentation Learning

Similar Papers 제목 키워드 기반

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

2025-06-16 · Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou 외

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatena…

Large Language Modelmultimodal interaction

Learning Unseen Modality Interaction

2023-06-22 · NeurIPS 2023 11

Multimodal learning assumes all modality combinations of interest are available during training to learn cross-modal correspondences. In this paper, we challenge this modality-complete assumption for multimodal learning …

RetrievalVideo Classification

Multimodal Graph Neural Network for Recommendation with Dynamic De-redundancy and Modality-Guided Feature De-noisy

2024-11-03 · Feng Mo, Lin Xiao, Qiya Song, Xieping Gao 외

Graph neural networks (GNNs) have become crucial in multimodal recommendation tasks because of their powerful ability to capture complex relationships between neighboring nodes. However, increasing the number of propagat…

DenoisingGraph Neural NetworkMultimodal Recommendation

Modality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation

2026-02-24 · Ji Dai, Quan Fang, Dengsheng Cai arxiv

Multimodal recommendation enhances ranking by integrating user-item interactions with item content, which is particularly effective under sparse feedback and long-tail distributions. However, multimodal signals are inher…

Multimodal RecommendationGraph Learning

Learning Multimodal Data Augmentation in Feature Space

2022-12-29 · Zichang Liu, Zhiqiang Tang, Xingjian Shi, Aston Zhang 외

The ability to jointly learn from multiple modalities, such as text, audio, and visual data, is a defining feature of intelligent systems. While there have been promising advances in designing neural networks to harness …

Data Augmentationimage-classificationImage ClassificationMultimodal Deep Learning