Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning
We present a simple approach which can turn a ViT encoder into an efficient video model, which can seamlessly work with both image and video inputs. By sparsely sampling the inputs, the model is able to do training and inference from both inputs. The model is easily scalable and can be adapted to large-scale pre-trained ViTs without requiring full finetuning. The model achieves SOTA results and the code will be open-sourced.
Code (1)
Tasks
Action ClassificationAction RecognitionAction Recognition In VideosSimilar Papers 제목 키워드 기반
Human Action Localization with Sparse Spatial Supervision
We introduce an approach for spatio-temporal human action localization using sparse spatial supervision. Our method leverages the large amount of annotated humans available today and extracts human tubes by combining a s…
Action LocalizationDiversityA Low-Computational Video Synopsis Framework with a Standard Dataset
Video synopsis is an efficient method for condensing surveillance videos. This technique begins with the detection and tracking of objects, followed by the creation of object tubes. These tubes consist of sequences, each…
Objectobject-detectionObject DetectionObject Tracking+1Spot On: Action Localization from Pointly-Supervised Proposals
We strive for spatio-temporal localization of actions in videos. The state-of-the-art relies on action proposals at test time and selects the best one with a classifier trained on carefully annotated box annotations. Ann…
Action LocalizationMultiple Instance LearningTemporal LocalizationPerson Re-identification in Videos by Analyzing Spatio-Temporal Tubes
Typical person re-identification frameworks search for k best matches in a gallery of images that are often collected in varying conditions. The gallery may contain image sequences when re-identification is done on video…
Person Re-IdentificationTemporal SequencesMotion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture tempo…
Contrastive LearningGraph GenerationPanoptic Scene Graph GenerationRelation+2