paper-with-me

홈 › Papers

MAFM^3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI

2025-11-14 · Mohammad Areeb Qazi, Munachiso S Nwadike, Ibrahim Almakky, Mohammad Yaqub, Numan Saeed arxiv

Foundational models are trained on extensive datasets to capture the general trends of a domain. However, in medical imaging, the scarcity of data makes pre-training for every domain, modality, or task challenging. Instead of building separate models, we propose MAFM^3 (Modular Adaptation of Foundation Models for Multi-Modal Medical AI), a framework that enables a single foundation model to expand into diverse domains, tasks, and modalities through lightweight modular components. These components serve as specialized skill sets that allow the system to flexibly activate the appropriate capability at the inference time, depending on the input type or clinical objective. Unlike conventional adaptation methods that treat each new task or modality in isolation, MAFM^3 provides a unified and expandable framework for efficient multitask and multimodality adaptation. Empirically, we validate our approach by adapting a chest CT foundation model initially trained for classification into prognosis and segmentation modules. Our results show improved performance on both tasks. Furthermore, by incorporating PET scans, MAFM^3 achieved an improvement in the Dice score 5% compared to the respective baselines. These findings establish that foundation models, when equipped with modular components, are not inherently constrained to their initial training scope but can evolve into multitask, multimodality systems for medical imaging. The code implementation of this work can be found at https://github.com/Areeb2735/CTscan_prognosis_VLM

📄 PDF Abstract BibTeX arXiv:2511.11212

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

2025-06-02 · Mohd Mujtaba Akhtar, Orchid Chetia Phukan, Girish, Swarup Ranjan Behera 외

In this work, we focus on non-verbal vocal sounds emotion recognition (NVER). We investigate mamba-based audio foundation models (MAFMs) for the first time for NVER and hypothesize that MAFMs will outperform attention-ba…

Emotion RecognitionMambaSpeech Emotion RecognitionSynthetic Speech Detection

Leveraging Generic Foundation Models for Multimodal Surgical Data Analysis

2025-09-08 · Simon Pezold, Jérôme A. Kurylec, Jan S. Liechti, Beat P. Müller 외 arxiv

We investigate how both the adaptation of a generic foundation model via transfer learning and the integration of complementary modalities from the operating room (OR) can support surgical data science. To this end, we u…

Surgical phase recognitionTransfer LearningDomain Adaptation

Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models

2025-01-30 · Hao Dong, Moru Liu, Kaiyang Zhou, Eleni Chatzi 외

In real-world scenarios, achieving domain adaptation and generalization poses significant challenges, as models must adapt to or generalize across unknown target distributions. Extending these capabilities to unseen mult…

Action RecognitionDomain AdaptationDomain GeneralizationSemantic Segmentation+1

Design Process of a Self Adaptive Smart Serious Games Ecosystem

2025-10-06 · X. Tao, P. Chen, M. Tsami, F. Khayati 외 arxiv

This paper outlines the design vision and planned evolution of Blexer v3, a modular and AI-driven rehabilitation ecosystem based on serious games. Building on insights from previous versions of the system, we propose a n…

Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

2026-05-21 · Yuting He, Chenyu You, Shuo Li arxiv

Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous imaging modalities. Monolithic self-supervised optimization on such dat…