paper-with-me

홈 › Papers

From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation

2025-11-27 · Zhen Chen, Yihang Fu, Gabriel Madera, Mauro Giuffre, Serina Applebaum, Hyunjae Kim, Hua Xu, Qingyu Chen arxiv

Multi-modal large language models (MLLMs) have shown promise in advancing healthcare. However, most existing models remain confined to single-image understanding, which greatly limits their applicability in clinical workflows. In practice, medical diagnosis and progression often require synthesizing information across multiple images from different modalities or time points. The development of medical MLLMs capable of such multi-image understanding has been hindered by the lack of large-scale, high-quality annotated training data. To address this limitation, we propose a novel framework that leverages license-permissive compound images in biomedical literature, as a rich yet underutilized data source for multi-image analysis. Specifically, we design a five-stage, context-aware instruction generation paradigm underpinned by a divide-and-conquer strategy. By decomposing multi-image analysis into manageable sub-tasks, this paradigm empowers MLLMs to move beyond single-panel analysis and provide a composite understanding by learning the complex spatial, temporal, and cross-modal relationships inherent in these compound figures. By parsing over 237,000 compound figures and their contextual text for instruction generation, we develop M3LLM, a medical multi-image multi-modal large language model. For benchmarking, we construct PMC-MI-Bench for composite understanding, manually validated by medical experts. Extensive experiments show that M3LLM significantly outperforms both general-purpose and specialized medical MLLMs across multi-image, single-image, text-only, and multi-choice scenarios. Notably, M3LLM exhibits strong generalization to longitudinal chest X-ray analysis using the MIMIC dataset. This work establishes a scalable and efficient paradigm for developing medical MLLMs capable of composite reasoning, bridging the gap between biomedical literature and real-world clinical applications.

📄 PDF Abstract BibTeX arXiv:2511.22232

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Diagnosis

Similar Papers 제목 키워드 기반

A Data Driven Approach for Compound Figure Separation Using Convolutional Neural Networks

2017-03-15 · Satoshi Tsutsui, David Crandall

A key problem in automatic analysis and understanding of scientific papers is to extract semantic information from non-textual paper components like figures, diagrams, tables, etc. Much of this work requires a very first…

Transfer Learning

A Two-stage Framework for Compound Figure Separation

2021-01-25 · Weixin Jiang, Eric Schwenker, Trevor Spreadbury, Nicola Ferrier 외

Scientific literature contains large volumes of complex, unstructured figures that are compound in nature (i.e. composed of multiple images, graphs, and drawings). Separation of these compound figures is critical for inf…

feature selectionInformation RetrievalRetrievalVocal Bursts Valence Prediction

Automatic Separation of Compound Figures in Scientific Articles

2016-06-03 · Mario Taschwer, Oge Marques

Content-based analysis and retrieval of digital images found in scientific articles is often hindered by images consisting of multiple subfigures (compound figures). We address this problem by proposing a method to autom…

ArticlesRetrieval

MedICaT: A Dataset of Medical Images, Captions, and Textual References

2020-10-12 · Findings of the Association for Computational Linguistics 2020 · Sanjay Subramanian, Lucy Lu Wang, Sachin Mehta, Ben Bogin 외

Understanding the relationship between figures and text is key to scientific document understanding. Medical figures in particular are quite complex, often consisting of several subfigures (75% of figures in our dataset)…

document understandingImage-text matchingRetrievalText Matching

Semantic Segmentation for Compound figures

2019-12-16 · Weixin Jiang, Eric Schwenker, Maria Chan, Oliver Cossairt

Scientific literature contains large volumes of unstructured data,with over 30\% of figures constructed as a combination of multiple images, these compound figures cannot be analyzed directly with existing information re…

Information RetrievalRetrievalSegmentationSemantic Segmentation