paper-with-me

Papers

GP-VLS: A general-purpose vision language model for surgery

2024-07-27 · Samuel Schmidgall, Joseph Cho, Cyril Zakka, William Hiesinger

Surgery requires comprehensive medical knowledge, visual assessment skills, and procedural expertise. While recent surgical AI models have focused on solving task-specific problems, there is a need for general-purpose systems that can understand surgical scenes and interact through natural language. This paper introduces GP-VLS, a general-purpose vision language model for surgery that integrates medical and surgical knowledge with visual scene understanding. For comprehensively evaluating general-purpose surgical models, we propose SurgiQual, which evaluates across medical and surgical knowledge benchmarks as well as surgical vision-language questions. To train GP-VLS, we develop six new datasets spanning medical knowledge, surgical textbooks, and vision-language pairs for tasks like phase recognition and tool identification. We show that GP-VLS significantly outperforms existing open- and closed-source models on surgical vision-language tasks, with 8-21% improvements in accuracy across SurgiQual benchmarks. GP-VLS also demonstrates strong performance on medical and surgical knowledge tests compared to open-source alternatives. Overall, GP-VLS provides an open-source foundation for developing AI assistants to support surgeons across a wide range of tasks and scenarios. The code and data for this work is publicly available at gpvls-surgery-vlm.github.io.

📄 PDF Abstract BibTeX arXiv:2407.19305

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingScene Understanding

Similar Papers 제목 키워드 기반

General-purpose foundation models for increased autonomy in robot-assisted surgery

2024-01-01 · Samuel Schmidgall, Ji Woong Kim, Alan Kuntz, Ahmed Ezzat Ghazi 외

The dominant paradigm for end-to-end robot learning focuses on optimizing task-specific objectives that solve a single robotic problem such as picking up an object or reaching a target position. However, recent work on h…

Vision-Language-Action

Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery

2025-03-29 · Boyi Ma, Yanguang Zhao, Jie Wang, Guankun Wang 외

The DeepSeek models have shown exceptional performance in general scene understanding, question-answering (QA), and text generation tasks, owing to their efficient training paradigm and strong reasoning capabilities. In …

Action UnderstandingInstrument RecognitionLarge Language ModelPosition+4

Benchmarking performance, explainability, and evaluation strategies of vision-language models for surgery: Challenges and opportunities

2025-05-16 · Jiajun Cheng, Xianwu Zhao, Shan Lin

Minimally invasive surgery (MIS) presents significant visual and technical challenges, including surgical instrument classification and understanding surgical action involving instruments, verbs, and anatomical targets. …

Benchmarking

Evaluating Large Vision-language Models for Surgical Tool Detection

2026-01-23 · Nakul Poudel, Richard Simon, Cristian A. Linte arxiv

Surgery is a highly complex process, and artificial intelligence has emerged as a transformative force in supporting surgical guidance and decision-making. However, the unimodal nature of most current AI systems limits t…

Zero-shot GeneralizationSurgical tool detectionInstrument Recognition

Imitation Learning for Robot Assistance in Open Surgery: A Multi-Policy Evaluation on Suture Following

2026-05-27 · Xucheng Wang, Zhizhou Yang, Xiaoman Zhang, Sung Eun Kim 외 arxiv

This study presents the first evaluation of general-purpose imitation learning for surgeon-robot collaborative assistance in open surgery, targeting suture following: the grab-pull-release motion an assistant performs at…