paper-with-me

Papers

CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection

2024-10-10 · Guankun Wang, Han Xiao, Huxin Gao, Renrui Zhang, Long Bai, Xiaoxiao Yang, Zhen Li, Hongsheng Li, Hongliang Ren

submucosal dissection (ESD) enables rapid resection of large lesions, minimizing recurrence rates and improving long-term overall survival. Despite these advantages, ESD is technically challenging and carries high risks of complications, necessitating skilled surgeons and precise instruments. Recent advancements in Large Visual-Language Models (LVLMs) offer promising decision support and predictive planning capabilities for robotic systems, which can augment the accuracy of ESD and reduce procedural risks. However, existing datasets for multi-level fine-grained ESD surgical motion understanding are scarce and lack detailed annotations. In this paper, we design a hierarchical decomposition of ESD motion granularity and introduce a multi-level surgical motion dataset (CoPESD) for training LVLMs as the robotic \textbf{Co}-\textbf{P}ilot of \textbf{E}ndoscopic \textbf{S}ubmucosal \textbf{D}issection. CoPESD includes 17,679 images with 32,699 bounding boxes and 88,395 multi-level motions, from over 35 hours of ESD videos for both robot-assisted and conventional surgeries. CoPESD enables granular analysis of ESD motions, focusing on the complex task of submucosal dissection. Extensive experiments on the LVLMs demonstrate the effectiveness of CoPESD in training LVLMs to predict following surgical robotic motions. As the first multimodal ESD motion dataset, CoPESD supports advanced research in ESD instruction-following and surgical automation. The dataset is available at \href{https://github.com/gkw0010/CoPESD}{https://github.com/gkw0010/CoPESD.}}

📄 PDF Abstract BibTeX arXiv:2410.07540

Code (1)

gkw0010/copesd 공식 구현 pytorch

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

2026-02-05 · Jinlin Wu, Felix Holm, Chuxi Chen, An Wang 외 arxiv

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level visual details, such as smoke, specular…

Action Triplet RecognitionPolyp SegmentationDepth Estimation

Joint Surgical Gesture and Task Classification with Multi-Task and Multimodal Learning

2018-05-02 · Duygu Sarikaya, Khurshid A. Guru, Jason J. Corso

We propose a novel multi-modal and multi-task architecture for simultaneous low level gesture and surgical task classification in Robot Assisted Surgery (RAS) videos.Our end-to-end architecture is based on the principles…

General ClassificationMulti-Task Learning

Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion

2026-01-18 · Meng Wei, Kun Yuan, Shi Li, Yue Zhou 외 arxiv

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizi…

Video Segmentation

TAToo: Vision-based Joint Tracking of Anatomy and Tool for Skull-base Surgery

2022-12-29 · Zhaoshuo Li, Hongchao Shu, Ruixing Liang, Anna Goodridge 외

Purpose: Tracking the 3D motion of the surgical tool and the patient anatomy is a fundamental requirement for computer-assisted skull-base surgery. The estimated motion can be used both for intra-operative guidance and f…

Anatomy

Recurrent and Spiking Modeling of Sparse Surgical Kinematics

2020-05-12 · Neil Getty, Zixuan Zhao, Stephan Gruessner, Liaohai Chen 외

Robot-assisted minimally invasive surgery is improving surgeon performance and patient outcomes. This innovation is also turning what has been a subjective practice into motion sequences that can be precisely measured. A…

BIG-bench Machine Learning