paper-with-me

홈 › Papers

ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy

2024-05-04 · Soumyadeep Chandra, Sayeed Shafayet Chowdhury, Courtney Yong, Chandru P. Sundaram, Kaushik Roy

Surgical action localization is a challenging computer vision problem. While it has promising applications including automated training of surgery procedures, surgical workflow optimization, etc., appropriate model design is pivotal to accomplishing this task. Moreover, the lack of suitable medical datasets adds an additional layer of complexity. To that effect, we introduce a new complex dataset of nephrectomy surgeries called UroSlice. To perform the action localization from these videos, we propose a novel model termed as `ViTALS' (Vision Transformer for Action Localization in Surgical Nephrectomy). Our model incorporates hierarchical dilated temporal convolution layers and inter-layer residual connections to capture the temporal correlations at finer as well as coarser granularities. The proposed approach achieves state-of-the-art performance on Cholec80 and UroSlice datasets (89.8% and 66.1% accuracy, respectively), validating its effectiveness.

📄 PDF Abstract BibTeX arXiv:2405.02571

Code (0)

등록된 구현이 없습니다.

Tasks

Action Localization

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery

2023-05-19 · Long Bai, Mobarakol Islam, Lalithkumar Seenivasan, Hongliang Ren

Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, expert surgeons are often overloaded with…

Answer Generationobject-detectionObject DetectionQuestion Answering+1

Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes

2025-10-28 · Jonas Hein, Lazaros Vlachopoulos, Maurits Geert Laurent Olthof, Bastian Sigrist 외 arxiv

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual…

Sound Source LocalizationScene UnderstandingPoint Clouds

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

2026-06-30 · Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado López 외 arxiv

Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation outputs do not explicitly encode the action-conditioned semantics that def…

Weakly Supervised YOLO Network for Surgical Instrument Localization in Endoscopic Videos

2023-09-23 · Rongfeng Wei, Jinlin Wu, Xuexue Bai, Ming Feng 외

In minimally invasive surgery, surgical instrument localization is a crucial task for endoscopic videos, which enables various applications for improving surgical outcomes. However, annotating the instrument localization…

Surgical Action Triplet Detection by Mixed Supervised Learning of Instrument-Tissue Interactions

2023-07-18 · Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter, Nicolas Padoy

Surgical action triplets describe instrument-tissue interactions as (instrument, verb, target) combinations, thereby supporting a detailed analysis of surgical scene activities and workflow. This work focuses on surgical…

Action Triplet DetectionTriplet