paper-with-me

Papers

LABRAD-OR: Lightweight Memory Scene Graphs for Accurate Bimodal Reasoning in Dynamic Operating Rooms

2023-03-23 · Ege Özsoy, Tobias Czempiel, Felix Holm, Chantal Pellegrini, Nassir Navab

Modern surgeries are performed in complex and dynamic settings, including ever-changing interactions between medical staff, patients, and equipment. The holistic modeling of the operating room (OR) is, therefore, a challenging but essential task, with the potential to optimize the performance of surgical teams and aid in developing new surgical technologies to improve patient outcomes. The holistic representation of surgical scenes as semantic scene graphs (SGG), where entities are represented as nodes and relations between them as edges, is a promising direction for fine-grained semantic OR understanding. We propose, for the first time, the use of temporal information for more accurate and consistent holistic OR modeling. Specifically, we introduce memory scene graphs, where the scene graphs of previous time steps act as the temporal representation guiding the current prediction. We design an end-to-end architecture that intelligently fuses the temporal information of our lightweight memory scene graphs with the visual information from point clouds and images. We evaluate our method on the 4D-OR dataset and demonstrate that integrating temporality leads to more accurate and consistent results achieving an +5% increase and a new SOTA of 0.88 in macro F1. This work opens the path for representing the entire surgery history with memory scene graphs and improves the holistic understanding in the OR. Introducing scene graphs as memory representations can offer a valuable tool for many temporal understanding tasks.

📄 PDF Abstract BibTeX arXiv:2303.13293

Code (1)

egeozsoy/LABRAD-OR 공식 구현 pytorch

Tasks

Scene Graph Generation

Similar Papers 제목 키워드 기반

Labrador: Exploring the Limits of Masked Language Modeling for Laboratory Data

2023-12-09 · David R. Bellamy, Bhawesh Kumar, Cindy Wang, Andrew Beam

In this work we introduce Labrador, a pre-trained Transformer model for laboratory data. Labrador and BERT were pre-trained on a corpus of 100 million lab test results from electronic health records (EHRs) and evaluated …

Language ModelingLanguage ModellingMasked Language ModelingTransfer Learning

PromptGCN: Bridging Subgraph Gaps in Lightweight GCNs

2024-10-14 · Shengwei Ji, Yujie Tian, Fei Liu, Xinlu Li 외

Graph Convolutional Networks (GCNs) are widely used in graph-based applications, such as social networks and recommendation systems. Nevertheless, large-scale graphs or deep aggregation layers in full-batch GCNs consume …

GPURecommendation Systems

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models

2026-05-13 · Vladislav Makarov, Mark Gizetdinov, Dmitry Yudin arxiv

Scene graph generation provides a compact structured representation for visual perception, but accurate and fast graph prediction from images and videos remains challenging. Recent VLM-based methods can generate scene gr…

Video scene graph generationReinforcement Learning

Genomic and pathological analyses of an asymmetric true hermaphroditism case in a female labrador retriever

2025-01-04 · Yihang Zhou

The two main gonadal development disorders in dogs are true hermaphroditism and XX male syndrome. True hermaphroditism can be divided into two subcategories: XX sex reversal and XY sex reversal. XX Sry-negative sex rever…

Memory Over Maps: 3D Object Localization Without Reconstruction

2026-03-20 · Rui Zhou, Xander Yap, Jianwen Cao, Allison Lau 외 arxiv

Target localization is a prerequisite for embodied tasks such as navigation and manipulation. Conventional approaches rely on constructing explicit 3D scene representations to enable target localization, such as point cl…

Object Localization3D ReconstructionRobot NavigationPoint Clouds