paper-with-me

Papers

Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning

2024-05-27 · Divyam Madaan, Taro Makino, Sumit Chopra, Kyunghyun Cho

Supervised multi-modal learning involves mapping multiple modalities to a target label. Previous studies in this field have concentrated on capturing in isolation either the inter-modality dependencies (the relationships between different modalities and the label) or the intra-modality dependencies (the relationships within a single modality and the label). We argue that these conventional approaches that rely solely on either inter- or intra-modality dependencies may not be optimal in general. We view the multi-modal learning problem from the lens of generative models where we consider the target as a source of multiple modalities and the interaction between them. Towards that end, we propose inter- & intra-modality modeling (I2M2) framework, which captures and integrates both the inter- and intra-modality dependencies, leading to more accurate predictions. We evaluate our approach using real-world healthcare and vision-and-language datasets with state-of-the-art models, demonstrating superior performance over traditional methods focusing only on one type of modality dependency.

📄 PDF Abstract BibTeX arXiv:2405.17613

Code (1)

divyam3897/i2m2 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Speaker-Guided Encoder-Decoder Framework for Emotion Recognition in Conversation

2022-06-07 · Yinan Bao, Qianwen Ma, Lingwei Wei, Wei Zhou 외

The emotion recognition in conversation (ERC) task aims to predict the emotion label of an utterance in a conversation. Since the dependencies between speakers are complex and dynamic, which consist of intra- and inter-s…

DecoderEmotion RecognitionEmotion Recognition in Conversation

Jointly Modeling Intra- and Inter-transaction Dependencies with Hierarchical Attentive Transaction Embeddings for Next-item Recommendation

2020-05-30 · Shoujin Wang, Longbing Cao, Liang Hu, Shlomo Berkovsky 외

A transaction-based recommender system (TBRS) aims to predict the next item by modeling dependencies in transactional data. Generally, two kinds of dependencies considered are intra-transaction dependency and inter-trans…

Recommendation Systems

Multi-Modality Cross Attention Network for Image and Sentence Matching

2020-06-01 · CVPR 2020 6 · Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang 외

The key of image and sentence matching is to accurately measure the visual-semantic similarity between an image and a sentence. However, most existing methods make use of only the intra-modality relationship within each …

Semantic SimilaritySemantic Textual SimilaritySentence

Probing Visual-Audio Representation for Video Highlight Detection via Hard-Pairs Guided Contrastive Learning

2022-06-21 · Shuaicheng Li, Feng Zhang, Kunlin Yang, Lingbo Liu 외

Video highlight detection is a crucial yet challenging problem that aims to identify the interesting moments in untrimmed videos. The key to this task lies in effective video representations that jointly pursue two goals…

Contrastive LearningHighlight DetectionRepresentation Learning

CG-DMER: Hybrid Contrastive-Generative Framework for Disentangled Multimodal ECG Representation Learning

2026-02-24 · Ziwei Niu, Hao Sun, Shujun Bian, Xihong Yang 외 arxiv

Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but…

Representation Learning