paper-with-me

홈 › Papers

Training Medical Large Vision-Language Models with Abnormal-Aware Feedback

2025-01-02 · Yucheng Zhou, Lingran Song, Jianbing Shen

Existing Medical Large Vision-Language Models (Med-LVLMs), which encapsulate extensive medical knowledge, demonstrate excellent capabilities in understanding medical images and responding to human queries based on these images. However, there remain challenges in visual localization in medical images, which is crucial for abnormality detection and interpretation. To address these issues, we propose a novel UMed-LVLM designed with Unveiling Medical abnormalities. Specifically, we collect a Medical Abnormalities Unveiling (MAU) dataset and propose a two-stage training method for UMed-LVLM training. To collect MAU dataset, we propose a prompt method utilizing the GPT-4V to generate diagnoses based on identified abnormal areas in medical images. Moreover, the two-stage training method includes Abnormal-Aware Instruction Tuning and Abnormal-Aware Rewarding, comprising Abnormal Localization Rewarding and Vision Relevance Rewarding. Experimental results demonstrate that our UMed-LVLM surpasses existing Med-LVLMs in identifying and understanding medical abnormality. In addition, this work shows that enhancing the abnormality detection capabilities of Med-LVLMs significantly improves their understanding of medical images and generalization capability.

📄 PDF Abstract BibTeX arXiv:2501.01377

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly DetectionVisual Localization

Similar Papers 제목 키워드 기반

3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection

2026-02-27 · Haowen Zhu, Ning Yin, Xiaogen Zhou arxiv

Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical imaging introduces two principal challenges: (1) modality-specific vision…

Representation Learning

Medical Vision-Language Pre-Training for Brain Abnormalities

2024-04-27 · Masoud Monajatipoor, Zi-Yi Dou, Aichi Chien, Nanyun Peng 외

Vision-language models have become increasingly powerful for tasks that require an understanding of both visual and linguistic elements, bridging the gap between these modalities. In the context of multimodal clinical AI…

Language ModelingLanguage Modelling

Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions

2025-03-05 · Jun Li, Che Liu, Wenjia Bai, Rossella Arcucci 외

Visual Language Models (VLMs) have demonstrated impressive capabilities in visual grounding tasks. However, their effectiveness in the medical domain, particularly for abnormality detection and localization within medica…

Anomaly DetectionVisual Grounding

Comprehensive language-image pre-training for 3D medical image understanding

2025-10-16 · Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor 외 arxiv

Vision-language pre-training, i.e., aligning images with paired text, is a powerful paradigm to create encoders that can be directly used for tasks such as classification, retrieval, and segmentation. In the 3D medical i…

Semantic Segmentation

Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis

2026-02-01 · Haoran Lai, Zihang Jiang, Kun Zhang, Qingsong Yao 외 arxiv

Developing 3D vision-language models with robust clinical reasoning remains a challenge due to the inherent complexity of volumetric medical imaging, the tendency of models to overfit superficial report patterns, and the…

Visual Question AnsweringReinforcement Learning