paper-with-me

Papers

On the Compositional Generalization of Multimodal LLMs for Medical Imaging

2024-12-28 · Zhenyang Cai, Junying Chen, Rongsheng Wang, Weihong Wang, Yonglin Deng, Dingjie Song, Yize Chen, Zixu Zhang, Benyou Wang

Multimodal large language models (MLLMs) hold significant potential in the medical field, but their capabilities are often limited by insufficient data in certain medical domains, highlighting the need for understanding what kinds of images can be used by MLLMs for generalization. Current research suggests that multi-task training outperforms single-task as different tasks can benefit each other, but they often overlook the internal relationships within these tasks, providing limited guidance on selecting datasets to enhance specific tasks. To analyze this phenomenon, we attempted to employ compositional generalization (CG)-the ability of models to understand novel combinations by recombining learned elements-as a guiding framework. Since medical images can be precisely defined by Modality, Anatomical area, and Task, naturally providing an environment for exploring CG. Therefore, we assembled 106 medical datasets to create Med-MAT for comprehensive experiments. The experiments confirmed that MLLMs can use CG to understand unseen medical images and identified CG as one of the main drivers of the generalization observed in multi-task training. Additionally, further studies demonstrated that CG effectively supports datasets with limited data and delivers consistent performance across different backbones, highlighting its versatility and broad applicability. Med-MAT is publicly available at https://github.com/FreedomIntelligence/Med-MAT.

📄 PDF Abstract BibTeX arXiv:2412.20070

Code (1)

freedomintelligence/med-mat 공식 구현

Similar Papers 제목 키워드 기반

CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging

2025-11-14 · Pooja Singh, Siddhant Ujjain, Tapan Kumar Gandhi, Sandeep Kumar arxiv

Recent advances in multimodal large language models have enabled unified processing of visual and textual inputs, offering promising applications in general-purpose medical AI. However, their ability to generalize compos…

Visual Question Answering

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

2025-11-30 · Haozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu 외 arxiv

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CM…

Visual Question AnsweringMultimodal ReasoningObject Detection

Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation

2025-03-10 · Zhi Qin, Qianhui Gui, Mouxiao Bian, Rui Wang 외

Medical imaging quality control (QC) is essential for accurate diagnosis, yet traditional QC methods remain labor-intensive and subjective. To address this challenge, in this study, we establish a standardized dataset an…

Image Quality Assessment

OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks

2025-11-02 · Zhihao Peng, Cheng Wang, Shengyuan Liu, Zhiying Liang 외 arxiv

Brain imaging analysis is crucial for diagnosing and treating brain disorders, and multimodal large language models (MLLMs) are increasingly supporting it. However, current brain imaging visual question-answering (VQA) b…

OrthoDoc: Multimodal Large Language Model for Assisting Diagnosis in Computed Tomography

2024-08-30 · Youzhu Jin, Yichen Zhang

Multimodal large language models (MLLMs) have achieved significant success in the general field of image processing. Their emerging task generalization and freeform conversational capabilities can greatly facilitate medi…

Computed Tomography (CT)DiagnosticLanguage ModelingLanguage Modelling+4