Joint-task Self-supervised Learning for Temporal Correspondence
This paper proposes to learn reliable dense correspondence from videos in a self-supervised manner. Our learning process integrates two highly related tasks: tracking large image regions \emph{and} establishing fine-grained pixel-level associations between consecutive video frames. We exploit the synergy between both tasks through a shared inter-frame affinity matrix, which simultaneously models transitions between video frames at both the region- and pixel-levels. While region-level localization helps reduce ambiguities in fine-grained matching by narrowing down search regions; fine-grained matching provides bottom-up features to facilitate region-level localization. Our method outperforms the state-of-the-art self-supervised methods on a variety of visual correspondence tasks, including video-object and part-segmentation propagation, keypoint tracking, and object tracking. Our self-supervised method even surpasses the fully-supervised affinity feature representation obtained from a ResNet-18 pre-trained on the ImageNet.
Code (2)
Tasks
Object TrackingSelf-Supervised LearningSemi-Supervised Video Object SegmentationUnsupervised Video Object SegmentationSimilar Papers 제목 키워드 기반
Bridging Stereo Matching and Optical Flow via Spatiotemporal Correspondence
Stereo matching and flow estimation are two essential tasks for scene understanding, spatially in 3D and temporally in motion. Existing approaches have been focused on the unsupervised setting due to the limited resource…
Optical Flow EstimationScene UnderstandingStereo MatchingStereo Matching HandSelf-Supervised Cross-View Correspondence with Predictive Cycle Consistency
Learning self-supervised visual correspondence is a long-studied task fundamental to visual understanding and human perception. However, existing correspondence methods largely focus on small image transformations, s…
ColorizationImitation LearningObject TrackingModelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence Learning
This paper presents a self-supervised method for learning reliable visual correspondence from unlabeled videos. We formulate the correspondence as finding paths in a joint space-time graph, where nodes are grid patches s…
Contrastive LearningRelationWhen the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
The past decade has witnessed notable achievements in self-supervised learning for video tasks. Recent efforts typically adopt the Masked Video Modeling (MVM) paradigm, leading to significant progress on multiple video t…
Representation LearningSelf-Supervised LearningSpatial-then-Temporal Self-Supervised Learning for Video Correspondence
In low-level video analyses, effective representations are important to derive the correspondences between video frames. These representations have been learned in a self-supervised fashion from unlabeled images or video…
Contrastive LearningSelf-Supervised Learning