paper-with-me

Papers

MultiModN- Multimodal, Multi-Task, Interpretable Modular Networks

2023-09-25 · Vinitra Swamy, Malika Satayeva, Jibril Frej, Thierry Bossy, Thijs Vogels, Martin Jaggi, Tanja Käser, Mary-Anne Hartley

Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potential of multiple data types to create a shared feature space with aligned semantic meaning across inputs of drastically varying sizes (i.e. images, text, sound). Most current MM architectures fuse these representations in parallel, which not only limits their interpretability but also creates a dependency on modality availability. We present MultiModN, a multimodal, modular network that fuses latent representations in a sequence of any number, combination, or type of modality while providing granular real-time predictive feedback on any number or combination of predictive tasks. MultiModN's composable pipeline is interpretable-by-design, as well as innately multi-task and robust to the fundamental issue of biased missingness. We perform four experiments on several benchmark MM datasets across 10 real-world tasks (predicting medical diagnoses, academic performance, and weather), and show that MultiModN's sequential MM fusion does not compromise performance compared with a baseline of parallel fusion. By simulating the challenging bias of missing not-at-random (MNAR), this work shows that, contrary to MultiModN, parallel fusion baselines erroneously learn MNAR and suffer catastrophic failure when faced with different patterns of MNAR at inference. To the best of our knowledge, this is the first inherently MNAR-resistant approach to MM modeling. In conclusion, MultiModN provides granular insights, robustness, and flexibility without compromising performance.

📄 PDF Abstract BibTeX arXiv:2309.14118

Code (1)

epfl-iglobalhealth/multimodn 공식 구현 pytorch

Similar Papers 제목 키워드 기반

MultiMoDN—Multimodal, Multi-Task, Interpretable Modular Networks

2023-09-21 · NeurIPS 2023 11

Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potential of multiple data types to create a sh…

A Human-Centric Approach to Explainable AI for Personalized Education

2025-05-28 · Vinitra Swamy

Deep neural networks form the backbone of artificial intelligence research, with potential to transform the human experience in areas ranging from autonomous driving to personal assistants, healthcare to education. Howev…

Autonomous DrivingMixture-of-Experts

Multi-modal land cover mapping of remote sensing images using pyramid attention and gated fusion networks

2021-11-06 · Qinghui Liu, Michael Kampffmeyer, Robert Jenssen, Arnt-Børre Salberg

Multi-modality data is becoming readily available in remote sensing (RS) and can provide complementary information about the Earth's surface. Effective fusion of multi-modal information is thus important for various appl…

Land Cover Classification

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey

2026-05-22 · Liangwei Nathan Zheng, Wei Emma Zhang, Olaf Maennel, Lin Yue 외 arxiv

Mixture-of-Experts (MoE) presents a naturally compatible and scalable framework for multimodal learning, demonstrating strong adaptability across diverse modalities and tasks. Despite its growing success, a comprehensive…

Embodied Multimodal Multitask Learning

2019-02-04 · Devendra Singh Chaplot, Lisa Lee, Ruslan Salakhutdinov, Devi Parikh 외

Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for different multimodal tasks, such as semantic goal navigation and…

Deep Reinforcement LearningDisentanglementEmbodied Question AnsweringQuestion Answering+3