paper-with-me

홈 › Papers

Surgical Instruction Generation with Transformers

2021-07-14 · Jinglu Zhang, Yinyu Nie, Jian Chang, Jian Jun Zhang

Automatic surgical instruction generation is a prerequisite towards intra-operative context-aware surgical assistance. However, generating instructions from surgical scenes is challenging, as it requires jointly understanding the surgical activity of current view and modelling relationships between visual information and textual description. Inspired by the neural machine translation and imaging captioning tasks in open domain, we introduce a transformer-backboned encoder-decoder network with self-critical reinforcement learning to generate instructions from surgical images. We evaluate the effectiveness of our method on DAISI dataset, which includes 290 procedures from various medical disciplines. Our approach outperforms the existing baseline over all caption evaluation metrics. The results demonstrate the benefits of the encoder-decoder structure backboned by transformer in handling multimodal context.

📄 PDF Abstract BibTeX arXiv:2107.06964

Code (1)

xumengyaamy/swinmlp_trancap pytorch

Tasks

DecoderMachine TranslationReinforcement Learning (RL)Translation

Similar Papers 제목 키워드 기반

Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model

2025-05-12 · Wei Li, Ming Hu, Guoan Wang, Lihao Liu 외

In ophthalmic surgery, developing an AI system capable of interpreting surgical videos and predicting subsequent operations requires numerous ophthalmic surgical videos with high-quality annotations, which are difficult …

Video Generation

LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning

2024-08-15 · Jiajie Li, Garrett Skinner, Gene Yang, Brian R Quaranto 외

Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal model…

Answer GenerationQuestion-Answer-GenerationQuestion AnsweringVideo Question Answering

Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models

2024-10-13 · Juseong Jin, Chang Wook Jeong

Conversation agents powered by large language models are revolutionizing the way we interact with visual data. Recently, large vision-language models (LVLMs) have been extensively studied for both images and videos. Howe…

Instruction FollowingQuestion AnsweringVisual Question Answering

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

2026-06-24 · Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin 외 arxiv

We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly availabl…

Question Answering

Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery

2025-09-18 · Ana Davila, Jacinto Colan, Yasuhisa Hasegawa arxiv

Effective human-robot collaboration in surgery is affected by the inherent ambiguity of verbal communication. This paper presents a framework for a robotic surgical assistant that interprets and disambiguates verbal inst…