paper-with-me

홈 › Papers

LoGSAM: Parameter-Efficient Cross-Modal Grounding for MRI Segmentation

2026-03-18 · Mohammad Robaitul Islam Bhuiyan, Sheethal Bhat, Melika Qahqaie arxiv

Precise localization and delineation of brain tumors using magnetic resonance imaging (MRI) are essential for planning therapy and guiding surgical decisions. To address this, we propose LoGSAM, a parameter-efficient, detection-driven framework that transforms radiologist dictation into text prompts for foundation-model-based localization and segmentation. Radiologist speech is first transcribed and translated using a pretrained Whisper ASR model, followed by negation-aware clinical NLP to extract tumor-specific textual prompts. These prompts guide text-conditioned tumor localization via a LoRA-adapted vision-language detection model, Grounding DINO (GDINO). The predicted bounding boxes are used as prompts for MedSAM to generate pixel-level tumor masks without any additional fine-tuning. On BRISC 2025, LoGSAM attains a Dice score of 80.32\%, reaching 98.6\% of a fully fine-tuned GDINO + MedSAM baseline while training fewer than 5\% of its parameters, indicating a favorable accuracy/parameter trade-off. In addition, we evaluate the full pipeline using German dictations from a board-certified radiologist on unseen MRI scans, achieving 91.7\% case-level class-extraction accuracy. These results highlight the feasibility of constructing a modular speech-to-segmentation pipeline from pretrained foundation models with minimal parameter updates.

📄 PDF Abstract BibTeX arXiv:2603.17576

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

F-LMM: Grounding Frozen Large Multimodal Models

2024-06-09 · CVPR 2025 1 · Size Wu, Sheng Jin, Wenwei Zhang, Lumin Xu 외

Endowing Large Multimodal Models (LMMs) with visual grounding capability can significantly enhance AIs' understanding of the visual world and their interaction with humans. However, existing methods typically fine-tune t…

General KnowledgeInstruction FollowingQuestion AnsweringReferring Expression+3

Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding

2026-05-08 · Yuan Yao, Qiushi Yang, Humen Zhong, Jiangning Wei 외 arxiv

Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit strong open-world visual grounding, but th…

Referring Expression SegmentationVisual Grounding

Progressive Prompt-Guided Cross-Modal Reasoning for Referring Image Segmentation

2026-03-30 · Jiachen Li, Hongyun Wang, Jinyu Xu, Wenbo Jiang 외 arxiv

Referring image segmentation aims to localize and segment a target object in an image based on a free-form referring expression. The core challenge lies in effectively bridging linguistic descriptions with object-level v…

Semantic SegmentationInstance SegmentationReferring ExpressionImage Segmentation

SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding

2024-07-03 · Weitai Kang, Gaowen Liu, Mubarak Shah, Yan Yan

Different from Object Detection, Visual Grounding deals with detecting a bounding box for each text-image pair. This one box for each text-image data provides sparse supervision signals. Although previous works achieve i…

object-detectionObject DetectionregressionSegmentation+1

Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities

2025-04-02 · Jing Liu, Wenxuan Wang, Yisi Zhang, Yepeng Tang 외

Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES methods primarily address object-level grounding, real-world scenarios …

DescriptiveLarge Language ModelMultimodal Large Language ModelObject+3