A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simple objective that can easily generalize all these methods to space-time. Our objective encourages temporally-persistent features in the same video, and in spite of its simplicity, it works surprisingly well across: (i) different unsupervised frameworks, (ii) pre-training datasets, (iii) downstream datasets, and (iv) backbone architectures. We draw a series of intriguing observations from this study, e.g., we discover that encouraging long-spanned persistency can be effective even if the timespan is 60 seconds. In addition to state-of-the-art results in multiple benchmarks, we report a few promising cases in which unsupervised pre-training can outperform its supervised counterpart. Code is made available at https://github.com/facebookresearch/SlowFast
Code (2)
Tasks
Representation LearningSelf-Supervised Action RecognitionUnsupervised Pre-trainingSimilar Papers 제목 키워드 기반
A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action Recognition
The goal of building a benchmark (suite of datasets) is to provide a unified protocol for fair evaluation and thus facilitate the evolution of a specific area. Nonetheless, we point out that existing protocols of action …
Action RecognitionDomain AdaptationRepresentation LearningSelf-Supervised Learning+2Scalable Spatiotemporal Graph Neural Networks
Neural forecasting of spatiotemporal time series drives both research and industrial innovation in several relevant application domains. Graph neural networks (GNNs) are often the core component of the forecasting archit…
Temporal SequencesTime SeriesTime Series AnalysisBridging Stereo Matching and Optical Flow via Spatiotemporal Correspondence
Stereo matching and flow estimation are two essential tasks for scene understanding, spatially in 3D and temporally in motion. Existing approaches have been focused on the unsupervised setting due to the limited resource…
Optical Flow EstimationScene UnderstandingStereo MatchingStereo Matching HandA Spatiotemporal Volumetric Interpolation Network for 4D Dynamic Medical Image
Dynamic medical imaging is usually limited in application due to the large radiation doses and longer image scanning and reconstruction times. Existing methods attempt to reduce the dynamic sequence by interpolating the …
AnatomyUnsupervised Video Anomaly Detection with Diffusion Models Conditioned on Compact Motion Representations
This paper aims to address the unsupervised video anomaly detection (VAD) problem, which involves classifying each frame in a video as normal or abnormal, without any access to labels. To accomplish this, the proposed me…
Anomaly DetectionVideo Anomaly Detection