paper-with-me

홈 › Papers

General surgery vision transformer: A video pre-trained foundation model for general surgery

2024-03-09 · Samuel Schmidgall, Ji Woong Kim, Jeffrey Jopling, Axel Krieger

The absence of openly accessible data and specialized foundation models is a major barrier for computational research in surgery. Toward this, (i) we open-source the largest dataset of general surgery videos to-date, consisting of 680 hours of surgical videos, including data from robotic and laparoscopic techniques across 28 procedures; (ii) we propose a technique for video pre-training a general surgery vision transformer (GSViT) on surgical videos based on forward video prediction that can run in real-time for surgical applications, toward which we open-source the code and weights of GSViT; (iii) we also release code and weights for procedure-specific fine-tuned versions of GSViT across 10 procedures; (iv) we demonstrate the performance of GSViT on the Cholec80 phase annotation task, displaying improved performance over state-of-the-art single frame predictors.

📄 PDF Abstract BibTeX arXiv:2403.05949

Code (1)

samuelschmidgall/gsvit 공식 구현 pytorch

Tasks

Video Prediction

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Multi-Head Attention 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…

Similar Papers 제목 키워드 기반

Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer

2025-09-29 · Mohsen Ghafoorian, Denis Korzhenkov, Amirhossein Habibian arxiv

Transformer-based video diffusion models (VDMs) deliver state-of-the-art video generation quality but are constrained by the quadratic cost of self-attention, making long sequences and high resolutions computationally ex…

Video Generation

Self-Supervised Video Desmoking for Laparoscopic Surgery

2024-03-17 · Renlong Wu, Zhilu Zhang, Shuohao Zhang, Longfei Gou 외

Due to the difficulty of collecting real paired data, most existing desmoking methods train the models by synthesizing smoke, generalizing poorly to real surgical scenarios. Although a few works have explored single-imag…

PitVis-2023 Challenge: Workflow Recognition in videos of Endoscopic Pituitary Surgery

2024-09-02 · Adrito Das, Danyal Z. Khan, Dimitrios Psychogyios, Yitong Zhang 외

The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery: including which surgical steps are…

Instrument Recognition

A vision-language model and platform for temporally mapping surgery from video

2026-03-23 · Dani Kiyasseh arxiv

Mapping surgery is fundamental to developing operative guidelines and enabling autonomous robotic surgery. Recent advances in artificial intelligence (AI) have shown promise in mapping the behaviour of surgeons from vide…

Computational Efficiency

A real-time spatiotemporal AI model analyzes skill in open surgical videos

2021-12-14 · Emmett D. Goodman, Krishna K. Patel, Yilun Zhang, William Locke 외

Open procedures represent the dominant form of surgery worldwide. Artificial intelligence (AI) has the potential to optimize surgical practice and improve patient outcomes, but efforts have focused primarily on minimally…