paper-with-me

홈 › Papers

Encoding Surgical Videos as Latent Spatiotemporal Graphs for Object and Anatomy-Driven Reasoning

2023-12-11 · Aditya Murali, Deepak Alapatt, Pietro Mascagni, Armine Vardazaryan, Alain Garcia, Nariaki Okamoto, Didier Mutter, Nicolas Padoy

Recently, spatiotemporal graphs have emerged as a concise and elegant manner of representing video clips in an object-centric fashion, and have shown to be useful for downstream tasks such as action recognition. In this work, we investigate the use of latent spatiotemporal graphs to represent a surgical video in terms of the constituent anatomical structures and tools and their evolving properties over time. To build the graphs, we first predict frame-wise graphs using a pre-trained model, then add temporal edges between nodes based on spatial coherence and visual and semantic similarity. Unlike previous approaches, we incorporate long-term temporal edges in our graphs to better model the evolution of the surgical scene and increase robustness to temporary occlusions. We also introduce a novel graph-editing module that incorporates prior knowledge and temporal coherence to correct errors in the graph, enabling improved downstream task performance. Using our graph representations, we evaluate two downstream tasks, critical view of safety prediction and surgical phase recognition, obtaining strong results that demonstrate the quality and flexibility of the learned representations. Code is available at github.com/CAMMA-public/SurgLatentGraph.

📄 PDF Abstract BibTeX arXiv:2312.06829

Code (1)

camma-public/surglatentgraph 공식 구현 pytorch

Tasks

Action RecognitionAnatomySemantic SimilaritySemantic Textual SimilaritySurgical phase recognition

Similar Papers 제목 키워드 기반

Dynamic Scene Graph Representation for Surgical Video

2023-09-25 · Felix Holm, Ghazal Ghazaei, Tobias Czempiel, Ege Özsoy 외

Surgical videos captured from microscopic or endoscopic imaging devices are rich but complex sources of information, depicting different tools and anatomical structures utilized during an extended amount of time. Despite…

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

2026-02-05 · Jinlin Wu, Felix Holm, Chuxi Chen, An Wang 외 arxiv

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level visual details, such as smoke, specular…

Action Triplet RecognitionPolyp SegmentationDepth Estimation

Mission Balance: Generating Under-represented Class Samples using Video Diffusion Models

2025-05-14 · Danush Kumar Venkatesh, Isabel Funke, Micha Pfeiffer, Fiona Kolbinger 외

Computer-assisted interventions can improve intra-operative guidance, particularly through deep learning methods that harness the spatiotemporal information in surgical videos. However, the severe data imbalance often fo…

Action Recognition

SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction

2024-07-29 · Çağhan Köksal, Ghazal Ghazaei, Felix Holm, Azade Farshad 외

Graph-based holistic scene representations facilitate surgical workflow understanding and have recently demonstrated significant success. However, this task is often hindered by the limited availability of densely annota…

DisentanglementGraph GenerationScene Graph Generation

SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation

2026-03-13 · Sampath Rapuri, Lalithkumar Seenivasan, Dominik Schneider, Roger Soberanis-Mukul 외 arxiv

A surgical world model capable of generating realistic surgical action videos with precise control over tool-tissue interactions can address fundamental challenges in surgical AI and simulation -- from data scarcity and …

Action RecognitionVideo Generation