paper-with-me

홈 › Papers

DM$^2$S$^2$: Deep Multi-Modal Sequence Sets with Hierarchical Modality Attention

2022-09-07 · Shunsuke Kitada, Yuki Iwazaki, Riku Togashi, Hitoshi Iyatomi

There is increasing interest in the use of multimodal data in various web applications, such as digital advertising and e-commerce. Typical methods for extracting important information from multimodal data rely on a mid-fusion architecture that combines the feature representations from multiple encoders. However, as the number of modalities increases, several potential problems with the mid-fusion model structure arise, such as an increase in the dimensionality of the concatenated multimodal features and missing modalities. To address these problems, we propose a new concept that considers multimodal inputs as a set of sequences, namely, deep multimodal sequence sets (DM$^2$S$^2$). Our set-aware concept consists of three components that capture the relationships among multiple modalities: (a) a BERT-based encoder to handle the inter- and intra-order of elements in the sequences, (b) intra-modality residual attention (IntraMRA) to capture the importance of the elements in a modality, and (c) inter-modality residual attention (InterMRA) to enhance the importance of elements with modality-level granularity further. Our concept exhibits performance that is comparable to or better than the previous set-aware models. Furthermore, we demonstrate that the visualization of the learned InterMRA and IntraMRA weights can provide an interpretation of the prediction results.

📄 PDF Abstract BibTeX arXiv:2209.03126

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAST: Multimodal Abstractive Summarization with Trimodal Hierarchical Attention

2020-10-15 · EMNLP (nlpbt) 2020 11 · Aman Khullar, Udit Arora

This paper presents MAST, a new model for Multimodal Abstractive Text Summarization that utilizes information from all three modalities -- text, audio and video -- in a multimodal video. Prior work on multimodal abstract…

Abstractive Text SummarizationMultimodal Abstractive Text SummarizationText Summarization

Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations

2024-07-06 · Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 외

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural l…

Hierarchical Multi-modal Transformer for Cross-modal Long Document Classification

2024-07-14 · Tengfei Liu, Yongli Hu, Junbin Gao, Yanfeng Sun 외

Long Document Classification (LDC) has gained significant attention recently. However, multi-modal data in long documents such as texts and images are not being effectively utilized. Prior studies in this area have attem…

Document ClassificationSentence

MHVAE: a Human-Inspired Deep Hierarchical Generative Model for Multimodal Representation Learning

2020-06-04 · Miguel Vasco, Francisco S. Melo, Ana Paiva

Humans are able to create rich representations of their external reality. Their internal representations allow for cross-modality inference, where available perceptions can induce the perceptual experience of missing inp…

Representation Learning

Seq2Seq2Sentiment: Multimodal Sequence to Sequence Models for Sentiment Analysis

2018-07-11 · WS 2018 7 · Hai Pham, Thomas Manzini, Paul Pu Liang, Barnabas Poczos

Multimodal machine learning is a core research area spanning the language, visual and acoustic modalities. The central challenge in multimodal learning involves learning representations that can process and relate inform…

Multimodal Sentiment AnalysisSentiment AnalysisTranslation