paper-with-me

Papers

SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement

2025-07-03 · Zeyu Lei, Hongyuan Yu, Jinlin Wu, Zhen Chen arxiv

Precise surgical interventions are vital to patient safety, and advanced enhancement algorithms have been developed to assist surgeons in decision-making. Despite significant progress, these algorithms are typically designed for single tasks in specific scenarios, limiting their effectiveness in complex real-world situations. To address this limitation, we propose SurgVisAgent, an end-to-end intelligent surgical vision agent built on multimodal large language models (MLLMs). SurgVisAgent dynamically identifies distortion categories and severity levels in endoscopic images, enabling it to perform a variety of enhancement tasks such as low-light enhancement, overexposure correction, motion blur elimination, and smoke removal. Specifically, to achieve superior surgical scenario understanding, we design a prior model that provides domain-specific knowledge. Additionally, through in-context few-shot learning and chain-of-thought (CoT) reasoning, SurgVisAgent delivers customized image enhancements tailored to a wide range of distortion types and severity levels, thereby addressing the diverse requirements of surgeons. Furthermore, we construct a comprehensive benchmark simulating real-world surgical distortions, on which extensive experiments demonstrate that SurgVisAgent surpasses traditional single-task models, highlighting its potential as a unified solution for surgical assistance.

📄 PDF Abstract BibTeX arXiv:2507.02252

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Learning

Similar Papers 제목 키워드 기반

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

2025-08-30 · Zhen Chen, Xingjian Luo, Kun Yuan, Jinlin Wu 외 arxiv

Surgical video understanding is crucial for facilitating Computer-Assisted Surgery (CAS) systems. Despite significant progress in existing studies, two major limitations persist, including inadequate visual content perce…

Video Reconstruction

VS-Assistant: Versatile Surgery Assistant on the Demand of Surgeons

2024-05-14 · Zhen Chen, Xingjian Luo, Jinlin Wu, Danny T. M. Chan 외

The surgical intervention is crucial to patient healthcare, and many studies have developed advanced algorithms to provide understanding and decision-making assistance for surgeons. Despite great progress, these algorith…

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

2025-06-03 · Zhitao Zeng, Zhu Zhuo, Xiaojun Jia, Erli Zhang 외

Foundation models have achieved transformative success across biomedical domains by enabling holistic understanding of multimodal data. However, their application in surgery remains underexplored. Surgical intelligence p…

Language ModelingLanguage Modelling

Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting

2026-02-11 · Rishikesh Bhyri, Brian R Quaranto, Philip J Seger, Kaity Tung 외 arxiv

Accurate counting of surgical instruments in Operating Rooms (OR) is a critical prerequisite for ensuring patient safety during surgery. Despite recent progress of large visual-language models and agentic AI, accurately …

Spatial ReasoningObject DetectionVisual Reasoning

Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes

2025-10-28 · Jonas Hein, Lazaros Vlachopoulos, Maurits Geert Laurent Olthof, Bastian Sigrist 외 arxiv

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual…

Sound Source LocalizationScene UnderstandingPoint Clouds