paper-with-me

홈 › Papers

Exploiting Temporal State Space Sharing for Video Semantic Segmentation

2025-03-26 · CVPR 2025 1 · Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Henghui Ding, Jing Yang, Ender Konukoglu, Xue Geng, Xudong Jiang

Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we introduce a Temporal Video State Space Sharing (TV3S) architecture to leverage Mamba state space models for temporal feature sharing. Our model features a selective gating mechanism that efficiently propagates relevant information across video frames, eliminating the need for a memory-heavy feature pool. By processing spatial patches independently and incorporating shifted operation, TV3S supports highly parallel computation in both training and inference stages, which reduces the delay in sequential state space processing and improves the scalability for long video sequences. Moreover, TV3S incorporates information from prior frames during inference, achieving long-range temporal coherence and superior adaptability to extended sequences. Evaluations on the VSPW and Cityscapes datasets reveal that our approach outperforms current state-of-the-art methods, establishing a new standard for VSS with consistent results across long video sequences. By achieving a good balance between accuracy and efficiency, TV3S shows a significant advancement in spatiotemporal modeling, paving the way for efficient video analysis. The code is publicly available at https://github.com/Ashesham/TV3S.git.

📄 PDF Abstract BibTeX arXiv:2503.20824

Code (1)

ashesham/tv3s 공식 구현 pytorch

Tasks

MambaSemantic SegmentationState Space ModelsVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…

Similar Papers 제목 키워드 기반

k-Space Deep Learning for Parallel MRI: Application to Time-Resolved MR Angiography

2018-06-03 · Eunju Cha, Eung Yeop Kim, Jong Chul Ye

Time-resolved angiography with interleaved stochastic trajectories (TWIST) has been widely used for dynamic contrast enhanced MRI (DCE-MRI). To achieve highly accelerated acquisitions, TWIST combines the periphery of the…

Deep Siamese Networks with Bayesian non-Parametrics for Video Object Tracking

2018-11-18 · Anthony D. Rhodes, Manan Goel

We present a novel algorithm utilizing a deep Siamese neural network as a general object similarity function in combination with a Bayesian optimization (BO) framework to encode spatio-temporal information for efficient …

Bayesian OptimizationObjectObject TrackingVideo Object Tracking

A novel efficient Multi-view traffic-related object detection framework

2023-02-23 · Kun Yang, Jing Liu, Dingkang Yang, Hanqi Wang 외

With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exp…

input filteringModel Selectionobject-detectionObject Detection

Learning by Aligning Videos in Time

2021-03-31 · CVPR 2021 1 · Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed 외

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level information. We leverage a novel combination of…

Representation LearningRetrievalVideo Alignment

MEGAN: Memory Enhanced Graph Attention Network for Space-Time Video Super-Resolution

2021-10-28 · Chenyu You, Lianyi Han, Aosong Feng, Ruihan Zhao 외

Space-time video super-resolution (STVSR) aims to construct a high space-time resolution video sequence from the corresponding low-frame-rate, low-resolution video sequence. Inspired by the recent success to consider spa…

Graph AttentionSpace-time Video Super-resolutionSuper-ResolutionVideo Super-Resolution