paper-with-me

Papers

Quantifying Modality Contributions via Disentangling Multimodal Representations

2025-11-22 · Padegal Amit, Omkar Mahesh Kashyap, Namitha Rayasam, Nidhi Shekhar, Surabhi Narayan arxiv

Quantifying modality contributions in multimodal models remains a challenge, as existing approaches conflate the notion of contribution itself. Prior work relies on accuracy-based approaches, interpreting performance drops after removing a modality as indicative of its influence. However, such outcome-driven metrics fail to distinguish whether a modality is inherently informative or whether its value arises only through interaction with other modalities. This distinction is particularly important in cross-attention architectures, where modalities influence each other's representations. In this work, we propose a framework based on Partial Information Decomposition (PID) that quantifies modality contributions by decomposing predictive information in internal embeddings into unique, redundant, and synergistic components. To enable scalable, inference-only analysis, we develop an algorithm based on the Iterative Proportional Fitting Procedure (IPFP) that computes layer and dataset-level contributions without retraining. This provides a principled, representation-level view of multimodal behavior, offering clearer and more interpretable insights than outcome-based metrics.

📄 PDF Abstract BibTeX arXiv:2511.19470

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

2025-05-29 · Sean Foley, Hong Nguyen, JIhwan Lee, Sudarsana Reddy Kadiri 외

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on…

Phoneme Recognition

Disentangling by Partitioning: A Representation Learning Framework for Multimodal Sensory Data

2018-05-29 · Wei-Ning Hsu, James Glass

Multimodal sensory data resembles the form of information perceived by humans for learning, and are easy to obtain in large quantities. Compared to unimodal data, synchronization of concepts between modalities in such da…

Representation LearningVariational Inference

Triple Disentangled Representation Learning for Multimodal Affective Analysis

2024-01-29 · Ying Zhou, Xuefeng Liang, Han Chen, Yin Zhao 외

Multimodal learning has exhibited a significant advantage in affective analysis tasks owing to the comprehensive information of various modalities, particularly the complementary information. Thus, many emerging studies …

DisentanglementRepresentation Learning

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

2026-07-29 · Vahidin Hasic, Chao Wang, Luis C. Garcia-Peraza-Herrera, David Watson 외 arxiv

Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regi…

Decision Making

Uncertainty-Aware Multimodal Learning via Conformal Shapley Intervals

2026-01-30 · Mathew Chandy, Michael Johnson, Judong Shen, Devan V. Mehrotra 외 arxiv

Multimodal learning combines information from multiple data modalities to improve predictive performance. However, modalities often contribute unequally and in a data dependent way, making it unclear which data modalitie…