Learning Long Range Spatio-Temporal Representations over Continuous Time Dynamic Graphs with State Space Models
Continuous-time dynamic graphs (CTDGs) provide a richer framework to capture fine-grained temporal patterns in evolving relational data. Long-range information propagation is a key challenge while learning representations, wherein it is important to retain and update information over long temporal horizons. Existing approaches restrict models to capture one-hop or local temporal neighborhoods and fail to capture multi-hop or global structural patterns. To mitigate this, we derive a parameter-efficient state-space modeling framework for continuous-time dynamic graphs (CTDG-SSM) from first principles. We first introduce continuous-time Topology-Aware higher order polynomial projection operator (CTT-HiPPO), a novel memory-based reformulation of HiPPO to jointly encode temporal dynamics and graph structure. The solution from CTT-HiPPO is obtained by projecting the classical HiPPO solution through a polynomial of the Laplacian matrix, yielding topology-aware memory updates that admit an equivalent state-space formulation for CTDGs (CTDG-SSM). Then a computationally efficient discrete formulation is obtained using the zero-order hold approach for model implementation. Across benchmarks on dynamic link prediction, dynamic node classification, and sequence classification, CTDG-SSM achieves state-of-the-art performance. Notably, it achieves large performance gains on datasets that require long range temporal (LRT) and spatial reasoning.
Code (0)
등록된 구현이 없습니다.
Tasks
Dynamic Link PredictionNode ClassificationSpatial ReasoningSimilar Papers 제목 키워드 기반
Facial Expression Analysis Using Decomposed Multiscale Spatiotemporal Networks
Video-based analysis of facial expressions has been increasingly applied to infer health states of individuals, such as depression and pain. Among the existing approaches, deep learning models composed of structures for …
Depression DetectionSpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
Human pose estimation in videos remains a challenge, largely due to the reliance on extensive manual annotation of large datasets, which is expensive and labor-intensive. Furthermore, existing approaches often struggle t…
Pose EstimationGLSFormer : Gated - Long, Short Sequence Transformer for Step Recognition in Surgical Videos
Automated surgical step recognition is an important task that can significantly improve patient safety and decision-making during surgeries. Existing state-of-the-art methods for surgical step recognition either rely on …
Decision MakingDemand Forecasting in Bike-sharing Systems Based on A Multiple Spatiotemporal Fusion Network
Bike-sharing systems (BSSs) have become increasingly popular around the globe and have attracted a wide range of research interests. In this paper, the demand forecasting problem in BSSs is studied. Spatial and temporal …
Demand ForecastingEnsemble LearningFeature ImportanceV4D: 4D Convolutional Neural Networks for Video-level Representation Learning
Most existing 3D CNN structures for video representation learning are clip-based methods, and do not consider video-level temporal evolution of spatio-temporal features. In this paper, we propose Video-level 4D Convoluti…
Representation LearningVideo Recognition