paper-with-me

홈 › Papers

BreastGPT: A Multimodal Large Language Model for the Full Spectrum of Breast Cancer Clinical Routine

2026-06-03 · Yang Liu, Jiajin Zhang, Danyang Tu, Yaojun Hu, Jiao Qu, Jiuyu Zhang, Yu Shi, Wei Fang, Shi Gu, Ling Zhang, Yingda Xia arxiv

Breast cancer remains a leading cause of cancer-related mortality among women. Its clinical management requires multimodal reasoning across a clinical workflow that spans \textit{screening}, \textit{diagnosis} and \textit{treatment planning}, where each stage involves distinct imaging modalities, task objectives, and reasoning patterns. However, constrained by data scarcity and model versatility, existing medical MLLMs are typically evaluated on isolated modalities or narrow task families, limiting their ability to support workflow-level clinical reasoning. In this work, we first introduce \textbf{BreastStage}, a workflow-aligned breast imaging instruction corpus comprising 1.86M instruction-following pairs curated from 17 sub-datasets across 5 imaging modalities and 136 task templates. Its held-out split, \textbf{BreastStage-Bench}, provides a comprehensive benchmark for evaluating multimodal reasoning across the breast cancer care continuum. Building on this corpus, we propose \textbf{BreastGPT}, a unified MLLM equipped with a dual-branch visual encoder and concept-preserving token compression to bridge the scale gap between standard radiology and gigapixel pathology. On BreastStage-Bench, BreastGPT achieves 75.66\% closed-ended accuracy and 89.92\% open-ended score, outperforming both general-purpose and medical-specific MLLMs across clinical stages and task formats. These results suggest that workflow-aligned data and cross-scale visual modeling are critical for clinically grounded medical MLLMs. All data, code, and model checkpoints are released at https://yangyy-liu.github.io/BreastGPT.io.

📄 PDF Abstract BibTeX arXiv:2606.04911

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

2024-10-21 · Xiang Yue, Yueqi Song, Akari Asai, Seungone Kim 외

Despite recent advances in multimodal large language models (MLLMs), their development has predominantly focused on English- and western-centric datasets and tasks, leaving most of the world's languages and diverse cultu…

SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

2024-02-08 · Dongyang Liu, Renrui Zhang, Longtian Qiu, Siyuan Huang 외

We propose SPHINX-X, an extensive Multimodality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX framework by removing redundant visual e…

BenchmarkingDiversityLanguage ModelingLanguage Modelling+4

Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models

2026-06-14 · Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen, Ke-Yue Zhang 외 arxiv

Multimodal large language models (MLLMs) have been increasingly adopted in forensics for their robust semantic understanding. As AI-generated images become realistic, semantic-level inconsistencies alone are often insuff…

Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization

2025-02-05 · Chanhui Lee, Hanbum Ko, Yuheon Song, Yongjun Jeong 외

Recent advances in large language models (LLMs) have led to models that tackle diverse molecular tasks, such as chemical reaction prediction and molecular property prediction. Large-scale molecular instruction-tuning dat…

Chemical Reaction PredictionMolecular Property PredictionMolecule CaptioningPrediction+1

MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework

2026-08-27 · Hai-tao Yu, Nan Min, Zheng Fang, Hongyu Zhan 외 arxiv

Inferring molecular structures from multimodal spectroscopic measurements requires integrating complementary yet highly heterogeneous signals. However, the common paradigm of directly concatenating multispectral sequence…