paper-with-me

홈 › Papers

Rad-VLSM: A Cross-Modal Framework with Semantics-Assisted Prompting for Medical Segmentation and Diagnosis

2026-05-18 · Fengyi Zhang, Xujie Zeng, Mohan Liu, Zengyi Wang, Yalong Jiang arxiv

Medical image segmentation is more clinically valuable when it supports diagnosis rather than merely producing lesion masks. However, diagnostically relevant lesion cues are often subtle and localized, while existing models may be distracted by background tissues, acoustic artifacts, and irrelevant visual correlations. To address this problem, we propose Rad-VLSM, a two-stage cross-modal framework for semantics-assisted lesion focusing, robust segmentation, and visually grounded diagnosis. In the first stage, a BLIP-2-based vision-language alignment module identifies lesion-related candidate regions under semantic guidance and converts them into box prompts. In the second stage, these prompts are fed into a SAM-based multitask network, where a multi-candidate region aggregation strategy improves prompt stability and guides lesion segmentation. The predicted masks are then used as spatial priors for diagnosis, and a visual-radiomics fusion head integrates lesion-aware visual features with selected radiomics descriptors. By using semantic information for localization rather than direct prediction, Rad-VLSM reduces text-to-diagnosis dependence and grounds diagnosis in lesion-level evidence. Experiments on a private clinical breast ultrasound dataset and public benchmarks show that Rad-VLSM achieves strong segmentation and diagnostic performance with favorable generalization.

📄 PDF Abstract BibTeX arXiv:2605.18130

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Image SegmentationLesion Segmentation

Similar Papers 제목 키워드 기반

Exploring Transfer Learning in Medical Image Segmentation using Vision-Language Models

2023-08-15 · Kanchan Poudel, Manish Dhakal, Prasiddha Bhandari, Rabin Adhikari 외

Medical image segmentation allows quantifying target structure size and shape, aiding in disease diagnosis, prognosis, surgery planning, and comprehension.Building upon recent advancements in foundation Vision-Language M…

Image SegmentationMedical Image SegmentationPrognosisSegmentation+3

Synthetic Boost: Leveraging Synthetic Data for Enhanced Vision-Language Segmentation in Echocardiography

2023-09-22 · Rabin Adhikari, Manish Dhakal, Safal Thapaliya, Kanchan Poudel 외

Accurate segmentation is essential for echocardiography-based assessment of cardiovascular diseases (CVDs). However, the variability among sonographers and the inherent challenges of ultrasound images hinder precise segm…

SegmentationVision-Language Segmentation

TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models

2024-10-07 · Rabin Adhikari, Safal Thapaliya, Manish Dhakal, Bishesh Khanal

Vision-Language Models (VLMs) have shown impressive performance in vision tasks, but adapting them to new domains often requires expensive fine-tuning. Prompt tuning techniques, including textual, visual, and multimodal …

BenchmarkingSegmentationVision-Language SegmentationVisual Prompt Tuning

Adversarial Robustness Analysis of Vision-Language Models in Medical Image Segmentation

2025-05-05 · Anjila Budathoki, Manish Dhakal

Adversarial attacks have been fairly explored for computer vision and vision-language models. However, the avenue of adversarial attack for the vision language segmentation models (VLSMs) is still under-explored, especia…

Adversarial AttackAdversarial RobustnessImage SegmentationMedical Image Analysis+3

VLSM-Adapter: Finetuning Vision-Language Segmentation Efficiently with Lightweight Blocks

2024-05-10 · Manish Dhakal, Rabin Adhikari, Safal Thapaliya, Bishesh Khanal

Foundation Vision-Language Models (VLMs) trained using large-scale open-domain images and text pairs have recently been adapted to develop Vision-Language Segmentation Models (VLSMs) that allow providing text prompts dur…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation+1