paper-with-me

Papers

Multimodal Language Analysis with Recurrent Multistage Fusion

2018-08-12 · EMNLP 2018 10 · Paul Pu Liang, Ziyin Liu, Amir Zadeh, Louis-Philippe Morency

Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities. Comprehending multimodal language requires modeling not only the interactions within each modality (intra-modal interactions) but more importantly the interactions between modalities (cross-modal interactions). In this paper, we propose the Recurrent Multistage Fusion Network (RMFN) which decomposes the fusion problem into multiple stages, each of them focused on a subset of multimodal signals for specialized, effective fusion. Cross-modal interactions are modeled using this multistage fusion approach which builds upon intermediate representations of previous stages. Temporal and intra-modal interactions are modeled by integrating our proposed fusion approach with a system of recurrent neural networks. The RMFN displays state-of-the-art performance in modeling human multimodal language across three public datasets relating to multimodal sentiment analysis, emotion recognition, and speaker traits recognition. We provide visualizations to show that each stage of fusion focuses on a different subset of multimodal signals, learning increasingly discriminative multimodal representations.

📄 PDF Abstract BibTeX arXiv:1808.03920

Code (1)

righ120/multimodal_nlp

Tasks

Emotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

Similar Papers 제목 키워드 기반

Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain Videos

2020-11-01 · EMNLP 2020 11 · Nayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang 외

Multimodal summarization for open-domain videos is an emerging task, aiming to generate a summary from multisource information (video, audio, transcript). Despite the success of recent multiencoder-decoder frameworks on …

Decoder

Multitask Learning and Multistage Fusion for Dimensional Audiovisual Emotion Recognition

2020-02-26 · ICASSP 2020 4 · Bagus Tris Atmaja, Masato Akagi

Due to its ability to accurately predict emotional state using multimodal features, audiovisual emotion recognition has recently gained more interest from researchers. This paper proposes two methods to predict emotional…

AttributeEmotion Recognition

Decision-Focused Forecasting: A Differentiable Multistage Optimisation Architecture

2024-05-23 · Egon Peršak, Miguel F. Anjos

Most decision-focused learning work has focused on single stage problems whereas many real-world decision problems are more appropriately modelled using multistage optimisation. In multistage problems contextual informat…

Decision MakingDecision Making Under Uncertainty

Analyzing Unaligned Multimodal Sequence via Graph Convolution and Graph Pooling Fusion

2020-11-27 · Sijie Mai, Songlong Xing, Jiaxuan He, Ying Zeng 외

In this paper, we study the task of multimodal sequence analysis which aims to draw inferences from visual, language and acoustic sequences. A majority of existing works generally focus on aligned fusion, mostly at word …

Multistage linguistic conditioning of convolutional layers for speech emotion recognition

2021-10-13 · Andreas Triantafyllopoulos, Uwe Reichel, Shuo Liu, Stephan Huber 외

In this contribution, we investigate the effectiveness of deep fusion of text and audio features for categorical and dimensional speech emotion recognition (SER). We propose a novel, multistage fusion method where the tw…

Emotion RecognitionSpeech Emotion Recognition