paper-with-me

Papers

MediSee: Reasoning-based Pixel-level Perception in Medical Images

2025-04-15 · Qinyue Tong, Ziqian Lu, Jun Liu, Yangming Zheng, Zheming Lu

Despite remarkable advancements in pixel-level medical image perception, existing methods are either limited to specific tasks or heavily rely on accurate bounding boxes or text labels as input prompts. However, the medical knowledge required for input is a huge obstacle for general public, which greatly reduces the universality of these methods. Compared with these domain-specialized auxiliary information, general users tend to rely on oral queries that require logical reasoning. In this paper, we introduce a novel medical vision task: Medical Reasoning Segmentation and Detection (MedSD), which aims to comprehend implicit queries about medical images and generate the corresponding segmentation mask and bounding box for the target object. To accomplish this task, we first introduce a Multi-perspective, Logic-driven Medical Reasoning Segmentation and Detection (MLMR-SD) dataset, which encompasses a substantial collection of medical entity targets along with their corresponding reasoning. Furthermore, we propose MediSee, an effective baseline model designed for medical reasoning segmentation and detection. The experimental results indicate that the proposed method can effectively address MedSD with implicit colloquial queries and outperform traditional medical referring segmentation methods.

📄 PDF Abstract BibTeX arXiv:2504.11008

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningReasoning SegmentationSegmentation

Similar Papers 제목 키워드 기반

MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision

2025-08-11 · Zhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing 외 arxiv

Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current…

Reinforcement Learning

OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

2024-06-27 · Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan 외

Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, lar…

DecoderSegmentationUniversal SegmentationVideo Understanding

ARIADNE: A Perception-Reasoning Synergy Framework for Trustworthy Coronary Angiography Analysis

2026-03-19 · Zhan Jin, Yu Luo, Yizhou Zhang, Ziyang Cui 외 arxiv

Conventional pixel-wise loss functions fail to enforce topological constraints in coronary vessel segmentation, producing fragmented vascular trees despite high pixel-level accuracy. We present ARIADNE, a two-stage frame…

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

2025-11-27 · Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao 외 arxiv

Despite significant strides in medical foundation models, the ultrasound domain lacks a comprehensive solution capable of bridging low-level Ultrasound Grounded Perception (e.g., segmentation, localization) and high-leve…

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

2026-01-14 · Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang 외 arxiv

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained vi…

Visual Question AnsweringInteractive SegmentationMultimodal ReasoningSpatial Reasoning