paper-with-me

홈 › Papers

DIME: Fine-grained Interpretations of Multimodal Models via Disentangled Local Explanations

2022-03-03 · Yiwei Lyu, Paul Pu Liang, Zihao Deng, Ruslan Salakhutdinov, Louis-Philippe Morency

The ability for a human to understand an Artificial Intelligence (AI) model's decision-making process is critical in enabling stakeholders to visualize model behavior, perform model debugging, promote trust in AI models, and assist in collaborative human-AI decision-making. As a result, the research fields of interpretable and explainable AI have gained traction within AI communities as well as interdisciplinary scientists seeking to apply AI in their subject areas. In this paper, we focus on advancing the state-of-the-art in interpreting multimodal models - a class of machine learning methods that tackle core challenges in representing and capturing interactions between heterogeneous data sources such as images, text, audio, and time-series data. Multimodal models have proliferated numerous real-world applications across healthcare, robotics, multimedia, affective computing, and human-computer interaction. By performing model disentanglement into unimodal contributions (UC) and multimodal interactions (MI), our proposed approach, DIME, enables accurate and fine-grained analysis of multimodal models while maintaining generality across arbitrary modalities, model architectures, and tasks. Through a comprehensive suite of experiments on both synthetic and real-world multimodal tasks, we show that DIME generates accurate disentangled explanations, helps users of multimodal models gain a deeper understanding of model behavior, and presents a step towards debugging and improving these models for real-world deployment. Code for our experiments can be found at https://github.com/lvyiwei1/DIME.

📄 PDF Abstract BibTeX arXiv:2203.02013

Code (1)

lvyiwei1/dime 공식 구현 pytorch

Tasks

Decision MakingDisentanglementTime SeriesTime Series Analysis

Methods 이 논문이 사용한 방법론

DIME 설명 없음

Similar Papers 제목 키워드 기반

Towards Learning Fine-Grained Disentangled Representations from Speech

2018-08-08 · Yuan Gong, Christian Poellabauer

Learning disentangled representations of high-dimensional data is currently an active research area. However, compared to the field of computer vision, less work has been done for speech processing. In this paper, we pro…

Representation LearningSpeech Representation Learning

CG-DMER: Hybrid Contrastive-Generative Framework for Disentangled Multimodal ECG Representation Learning

2026-02-24 · Ziwei Niu, Hao Sun, Shujun Bian, Xihong Yang 외 arxiv

Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but…

Representation Learning

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

2026-07-30 · Bohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen 외 arxiv

Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing d…

Image RetrievalVisual Dialog

Refining Multidimensional Video Reward Models via Disentangled Influence Functions

2026-05-27 · Muyao Wang, Zeke Xie, Hideki Nakayama arxiv

As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various axes. To address this, recent works have focused on developing Multidi…

Discovering Semantic Subdimensions through Disentangled Conceptual Representations

2025-08-29 · Yunhao Zhang, Shaonan Wang, Nan Lin, Xinyi Dong 외 arxiv

Understanding the core dimensions of conceptual semantics is fundamental to uncovering how meaning is organized in language and the brain. Existing approaches often rely on predefined semantic dimensions that offer only …