paper-with-me

홈 › Papers

Tri-modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms

2024-04-14 · Diandian Guo, Manxi Lin, Jialun Pei, He Tang, Yueming Jin, Pheng-Ann Heng

A comprehensive understanding of surgical scenes allows for monitoring of the surgical process, reducing the occurrence of accidents and enhancing efficiency for medical professionals. Semantic modeling within operating rooms, as a scene graph generation (SGG) task, is challenging since it involves consecutive recognition of subtle surgical actions over prolonged periods. To address this challenge, we propose a Tri-modal (i.e., images, point clouds, and language) confluence with Temporal dynamics framework, termed TriTemp-OR. Diverging from previous approaches that integrated temporal information via memory graphs, our method embraces two advantages: 1) we directly exploit bi-modal temporal information from the video streaming for hierarchical feature interaction, and 2) the prior knowledge from Large Language Models (LLMs) is embedded to alleviate the class-imbalance problem in the operating theatre. Specifically, our model performs temporal interactions across 2D frames and 3D point clouds, including a scale-adaptive multi-view temporal interaction (ViewTemp) and a geometric-temporal point aggregation (PointTemp). Furthermore, we transfer knowledge from the biomedical LLM, LLaVA-Med, to deepen the comprehension of intraoperative relations. The proposed TriTemp-OR enables the aggregation of tri-modal features through relation-aware unification to predict relations so as to generate scene graphs. Experimental results on the 4D-OR benchmark demonstrate the superior performance of our model for long-term OR streaming.

📄 PDF Abstract BibTeX arXiv:2404.09231

Code (0)

등록된 구현이 없습니다.

Tasks

Graph GenerationScene Graph Generation

Similar Papers 제목 키워드 기반

Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation

2025-09-06 · Xiaomeng Zhu, Changwei Wang, Haozhe Wang, Xinyu Liu 외 arxiv

A scene graph is a structured representation of objects and their spatio-temporal relationships in dynamic scenes. Scene Graph Anticipation (SGA) involves predicting future scene graphs from video clips, enabling applica…

Relational Reasoning

Preemptive Spatiotemporal Trajectory Adjustment for Heterogeneous Vehicles in Highway Merging Zones

2025-09-30 · Yuan Li, Xiaoxue Xu, Xiang Dong, Junfeng Hao 외 arxiv

Aiming at the problem of driver's perception lag and low utilization efficiency of space-time resources in expressway ramp confluence area, based on the preemptive spatiotemporal trajectory Adjustment system, from the pe…

Autonomous Driving

AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency

2025-10-30 · Piyushkumar Patel arxiv

Text to video generation has emerged as a critical frontier in generative artificial intelligence, yet existing approaches struggle with maintaining temporal consistency, compositional understanding, and fine grained con…

Scene UnderstandingVideo Generation

Aion: Towards Hierarchical 4D Scene Graphs with Temporal Flow Dynamics

2025-12-10 · Iacopo Catalano, Eduardo Montijano, Javier Civera, Julio A. Placed 외 arxiv

Autonomous navigation in dynamic environments requires spatial representations that capture both semantic structure and temporal evolution. 3D Scene Graphs (3DSGs) provide hierarchical multi-resolution abstractions that …

Exploiting Long-Term Dependencies for Generating Dynamic Scene Graphs

2021-12-18 · Shengyu Feng, Subarna Tripathi, Hesham Mostafa, Marcel Nassar 외

Dynamic scene graph generation from a video is challenging due to the temporal dynamics of the scene and the inherent temporal fluctuations of predictions. We hypothesize that capturing long-term temporal dependencies is…

Graph GenerationObjectScene Graph DetectionScene Graph Generation+1