paper-with-me

Papers

Information-Theoretic Decomposition for Multimodal Interaction Learning

2026-06-10 · Zequn Yang, Yake Wei, Haotian Ni, Zhihao Xu, Di Hu arxiv

Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the first systematic, information-theoretic analysis highlighting why learning these dynamic, sample-specific interactions is critical for effective multimodal learning. Our analysis further reveals deficits in conventional paradigms at learning these distinct interaction types: modality ensemble approaches struggle to capture synergy, while joint learning paradigms often under-utilize redundant information. This highlights the need for an approach that can adaptively learn from different interaction types on a per-sample basis. To this end, we propose Decomposition-based Multimodal Interaction Learning (DMIL), a novel paradigm that explicitly models and learns from sample-specific interactions. First, we design a variational decomposition architecture to isolate the constituent interaction components. Second, we employ a new learning strategy that leverages these explicit interaction components in a fine-tuning process to achieve comprehensive interaction learning. Extensive experiments across diverse tasks and architectures demonstrate that DMIL consistently achieves superior performance by adapting to holistic sample-specific interactions. Our framework is flexible and broadly applicable, establishing an interaction-centric paradigm for multimodal learning. The code is available at https://github.com/GeWu-Lab/DMIL.

📄 PDF Abstract BibTeX arXiv:2606.11614

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quantifying & Modeling Multimodal Interactions: An Information Decomposition Framework

2023-02-23 · NeurIPS 2023 11 · Paul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling 외

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advance…

Model Selection

SPICE: Synergy and Partial Information Based Curriculum Evolution

2026-06-15 · Ankush Pratap Singh, Houwei Cao, Yong Liu arxiv

Multimodal learning exploits complementary information across heterogeneous modalities. The informativeness of each modality can vary widely across samples and training stages. Existing multimodal curriculum learning str…

Multimodal Fusion Interactions: A Study of Human and Automatic Quantification

2023-06-07 · Paul Pu Liang, Yun Cheng, Ruslan Salakhutdinov, Louis-Philippe Morency

In order to perform multimodal fusion of heterogeneous signals, we need to understand their interactions: how each modality individually provides information useful for a task and how this information changes in the pres…

counterfactual

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition

2026-05-31 · Wanlong Fang, Tianle Zhang, Wen Tao, Alvin Chan arxiv

Understanding modality interaction in multimodal large language models (MLLMs) is central to reliable deployment. We introduce Partial Information Decomposition (PID) as a decision-level framework that separates unique, …

Multimodal Reasoning

MUTAN: Multimodal Tucker Fusion for Visual Question Answering

2017-05-18 · ICCV 2017 10 · Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, Nicolas Thome

Bilinear models provide an appealing framework for mixing and merging information in Visual Question Answering (VQA) tasks. They help to learn high level associations between question meaning and visual concepts in the i…

Visual Question AnsweringVisual Question Answering (VQA)