Learning a Spatio-Temporal Embedding for Video Instance Segmentation
We present a novel embedding approach for video instance segmentation. Our method learns a spatio-temporal embedding integrating cues from appearance, motion, and geometry; a 3D causal convolutional network models motion, and a monocular self-supervised depth loss models geometry. In this embedding space, video-pixels of the same instance are clustered together while being separated from other instances, to naturally track instances over time without any complex post-processing. Our network runs in real-time as our architecture is entirely causal - we do not incorporate information from future frames, contrary to previous methods. We show that our model can accurately track and segment instances, even with occlusions and missed detections, advancing the state-of-the-art on the KITTI Multi-Object and Tracking Dataset.
Code (1)
Tasks
Instance SegmentationSemantic SegmentationVideo Instance SegmentationSimilar Papers 제목 키워드 기반
STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos
Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by-detection paradigm and model a video clip as a sequence of images. Multiple networks are used to de…
Instance SegmentationSemantic SegmentationUnsupervised Video Object SegmentationVideo Instance SegmentationSTC: Spatio-Temporal Contrastive Learning for Video Instance Segmentation
Video Instance Segmentation (VIS) is a task that simultaneously requires classification, segmentation, and instance association in a video. Recent VIS approaches rely on sophisticated pipelines to achieve this goal, incl…
Contrastive LearningInstance SegmentationSegmentationSemantic Segmentation+1Towards Robust Video Instance Segmentation with Temporal-Aware Transformer
Most existing transformer based video instance segmentation methods extract per frame features independently, hence it is challenging to solve the appearance deformation problem. In this paper, we observe the temporal in…
DecoderInstance SegmentationSemantic SegmentationVideo Instance SegmentationA2VIS: Amodal-Aware Approach to Video Instance Segmentation
Handling occlusion remains a significant challenge for video instance-level tasks like Multiple Object Tracking (MOT) and Video Instance Segmentation (VIS). In this paper, we propose a novel framework, Amodal-Aware Video…
Instance SegmentationMultiple Object TrackingObjectObject Tracking+3Deformable VisTR: Spatio temporal deformable attention for video instance segmentation
Video instance segmentation (VIS) task requires classifying, segmenting, and tracking object instances over all frames in a video clip. Recently, VisTR has been proposed as end-to-end transformer-based VIS framework, whi…
GPUInstance SegmentationSemantic SegmentationVideo Instance Segmentation