paper-with-me

홈 › Papers

Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role Labeling

2023-08-09 · Yu Zhao, Hao Fei, Yixin Cao, Bobo Li, Meishan Zhang, Jianguo Wei, Min Zhang, Tat-Seng Chua

Video Semantic Role Labeling (VidSRL) aims to detect the salient events from given videos, by recognizing the predict-argument event structures and the interrelationships between events. While recent endeavors have put forth methods for VidSRL, they can be mostly subject to two key drawbacks, including the lack of fine-grained spatial scene perception and the insufficiently modeling of video temporality. Towards this end, this work explores a novel holistic spatio-temporal scene graph (namely HostSG) representation based on the existing dynamic scene graph structures, which well model both the fine-grained spatial semantics and temporal dynamics of videos for VidSRL. Built upon the HostSG, we present a nichetargeting VidSRL framework. A scene-event mapping mechanism is first designed to bridge the gap between the underlying scene structure and the high-level event semantic structure, resulting in an overall hierarchical scene-event (termed ICE) graph structure. We further perform iterative structure refinement to optimize the ICE graph, such that the overall structure representation can best coincide with end task demand. Finally, three subtask predictions of VidSRL are jointly decoded, where the end-to-end paradigm effectively avoids error propagation. On the benchmark dataset, our framework boosts significantly over the current best-performing model. Further analyses are shown for a better understanding of the advances of our methods.

📄 PDF Abstract BibTeX arXiv:2308.05081

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Role Labeling

Similar Papers 제목 키워드 기반

STEGNav: Spatio-Temporal Event Graph Reasoning for Multimodal Lifelong Object Navigation

2026-08-28 · Yang Chen, Zhenyu Huang, Wenbo Fu, Danyang Peng 외 arxiv

Multimodal lifelong navigation requires an agent to autonomously explore unseen environments while sequentially completing navigation tasks specified by object categories, language descriptions, or reference images. Exis…

SANGRIA: Surgical Video Scene Graph Optimization for Surgical Workflow Prediction

2024-07-29 · Çağhan Köksal, Ghazal Ghazaei, Felix Holm, Azade Farshad 외

Graph-based holistic scene representations facilitate surgical workflow understanding and have recently demonstrated significant success. However, this task is often hindered by the limited availability of densely annota…

DisentanglementGraph GenerationScene Graph Generation

SpOT: Spatiotemporal Modeling for 3D Object Tracking

2022-07-12 · Colton Stearns, Davis Rempe, Jie Li, Rares Ambrus 외

3D multi-object tracking aims to uniquely and consistently identify all mobile entities through time. Despite the rich spatiotemporal information available in this setting, current 3D tracking methods primarily rely on a…

3D Multi-Object Tracking3D Object TrackingMulti-Object TrackingObject+1

Spatiotemporal Event Graphs for Dynamic Scene Understanding

2023-12-11 · Salman Khan

Dynamic scene understanding is the ability of a computer system to interpret and make sense of the visual information present in a video of a real-world scene. In this thesis, we present a series of frameworks for dynami…

Action DetectionActivity DetectionAutonomous DrivingContinual Learning+3

Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision

2021-12-09 · CVPR 2022 1 · Liangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong 외

Modern self-supervised learning algorithms typically enforce persistency of instance representations across views. While being very effective on learning holistic image and video representations, such an objective become…

Action LocalizationAction RecognitionContrastive LearningObject Tracking+3