paper-with-me

홈 › Papers

Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks

2025-11-10 · Lingran Song, Yucheng Zhou, Jianbing Shen arxiv

Despite significant progress in pixel-level medical image analysis, existing medical image segmentation models rarely explore medical segmentation and diagnosis tasks jointly. However, it is crucial for patients that models can provide explainable diagnoses along with medical segmentation results. In this paper, we introduce a medical vision-language task named Medical Diagnosis Segmentation (MDS), which aims to understand clinical queries for medical images and generate the corresponding segmentation masks as well as diagnostic results. To facilitate this task, we first present the Multimodal Multi-disease Medical Diagnosis Segmentation (M3DS) dataset, containing diverse multimodal multi-disease medical images paired with their corresponding segmentation masks and diagnosis chain-of-thought, created via an automated diagnosis chain-of-thought generation pipeline. Moreover, we propose Sim4Seg, a novel framework that improves the performance of diagnosis segmentation by taking advantage of the Region-Aware Vision-Language Similarity to Mask (RVLS2M) module. To improve overall performance, we investigate a test-time scaling strategy for MDS tasks. Experimental results demonstrate that our method outperforms the baselines in both segmentation and diagnosis.

📄 PDF Abstract BibTeX arXiv:2511.06665

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Image SegmentationMedical Diagnosis

Similar Papers 제목 키워드 기반

MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis

2025-06-23 · Yuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue 외

Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current approaches often rely on single-modal data,…

DiagnosticLarge Language ModelMultimodal Large Language Model

Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification

2025-08-06 · Jianxun Yu, Ruiquan Ge, Zhipeng Wang, Cheng Yang 외 arxiv

The diagnosis of medical diseases faces challenges such as the misdiagnosis of small lesions. Deep learning, particularly multimodal approaches, has shown great potential in the field of medical disease diagnosis. Howeve…

Feature-based Transformer with Incomplete Multimodal Brain Images for Diagnosis of Neurodegenerative Diseases

2023-07-01 · Conference 2023 7 · Xingyu Gao, Feng Shi, Dinggang Shen & Manhua Liu

Benefiting from complementary information, multimodal brain imaging analysis has distinct advantages over single-modal methods for the diagnosis of neurodegenerative diseases such as Alzheimer’s disease. However, multi-m…

OphGLM: Training an Ophthalmology Large Language-and-Vision Assistant based on Instructions and Dialogue

2023-06-21 · Weihao Gao, Zhuo Deng, Zhiyuan Niu, Fuju Rong 외

Large multimodal language models (LMMs) have achieved significant success in general domains. However, due to the significant differences between medical images and text and general web content, the performance of LMMs i…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+1

Can GPT-4V(ision) Serve Medical Applications? Case Studies on GPT-4V for Multimodal Medical Diagnosis

2023-10-15 · Chaoyi Wu, Jiayu Lei, Qiaoyu Zheng, Weike Zhao 외

Driven by the large foundation models, the development of artificial intelligence has witnessed tremendous progress lately, leading to a surge of general interest from the public. In this study, we aim to assess the perf…

AnatomyComputed Tomography (CT)Decision MakingMedical Diagnosis