paper-with-me

Papers

Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!

2020-10-13 · EMNLP 2020 11 · Jack Hessel, Lillian Lee

Modeling expressive cross-modal interactions seems crucial in multimodal tasks, such as visual question answering. However, sometimes high-performing black-box algorithms turn out to be mostly exploiting unimodal signals in the data. We propose a new diagnostic tool, empirical multimodally-additive function projection (EMAP), for isolating whether or not cross-modal interactions improve performance for a given model on a given task. This function projection modifies model predictions so that cross-modal interactions are eliminated, isolating the additive, unimodal structure. For seven image+text classification tasks (on each of which we set new state-of-the-art benchmarks), we find that, in many cases, removing cross-modal interactions results in little to no performance degradation. Surprisingly, this holds even when expressive models, with capacity to consider interactions, otherwise outperform less expressive models; thus, performance improvements, even when present, often cannot be attributed to consideration of cross-modal feature interactions. We hence recommend that researchers in multimodal machine learning report the performance not only of unimodal baselines, but also the EMAP of their best-performing model.

📄 PDF Abstract BibTeX arXiv:2010.06572

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticImage-text ClassificationQuestion Answeringtext-classificationText ClassificationVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Multimodal Language Analysis with Recurrent Multistage Fusion

2018-08-12 · EMNLP 2018 10 · Paul Pu Liang, Ziyin Liu, Amir Zadeh, Louis-Philippe Morency

Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities. Comprehending multimodal language requires modeling n…

Emotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective

2025-09-02 · Shijie Wang, Li Zhang, Xinyan Liang, Yuhua Qian 외 arxiv

Multimodal learning typically utilizes multimodal joint loss to integrate different modalities and enhance model performance. However, this joint learning strategy can induce modality imbalance, where strong modalities o…

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

2026-06-04 · Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello arxiv

Transformer-based multimodal models rely on attention mechanisms to integrate information across heterogeneous modalities. Despite their success, existing multimodal attention formulations compute their scores through co…

Multimodal Fusion Interactions: A Study of Human and Automatic Quantification

2023-06-07 · Paul Pu Liang, Yun Cheng, Ruslan Salakhutdinov, Louis-Philippe Morency

In order to perform multimodal fusion of heterogeneous signals, we need to understand their interactions: how each modality individually provides information useful for a task and how this information changes in the pres…

counterfactual

MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens

2023-10-03 · Kaizhi Zheng, Xuehai He, Xin Eric Wang

The effectiveness of Multimodal Large Language Models (MLLMs) demonstrates a profound capability in multimodal understanding. However, the simultaneous generation of images with coherent texts is still underdeveloped. Ad…

Image Generationmultimodal generationReading ComprehensionText Generation