paper-with-me

홈 › Papers

Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning

2022-12-06 · CVPR 2023 1 · AJ Piergiovanni, Weicheng Kuo, Anelia Angelova

We present a simple approach which can turn a ViT encoder into an efficient video model, which can seamlessly work with both image and video inputs. By sparsely sampling the inputs, the model is able to do training and inference from both inputs. The model is easily scalable and can be adapted to large-scale pre-trained ViTs without requiring full finetuning. The model achieves SOTA results and the code will be open-sourced.

📄 PDF Abstract BibTeX arXiv:2212.03229

Code (1)

daniel-code/TubeViT pytorch

Tasks

Action ClassificationAction RecognitionAction Recognition In Videos

Similar Papers 제목 키워드 기반

Human Action Localization with Sparse Spatial Supervision

2016-05-17 · Philippe Weinzaepfel, Xavier Martin, Cordelia Schmid

We introduce an approach for spatio-temporal human action localization using sparse spatial supervision. Our method leverages the large amount of annotated humans available today and extracts human tubes by combining a s…

Action LocalizationDiversity

A Low-Computational Video Synopsis Framework with a Standard Dataset

2024-09-08 · Ramtin Malekpour, M. Mehrdad Morsali, Hoda Mohammadzade

Video synopsis is an efficient method for condensing surveillance videos. This technique begins with the detection and tracking of objects, followed by the creation of object tubes. These tubes consist of sequences, each…

Objectobject-detectionObject DetectionObject Tracking+1

Spot On: Action Localization from Pointly-Supervised Proposals

2016-04-26 · Pascal Mettes, Jan C. van Gemert, Cees G. M. Snoek

We strive for spatio-temporal localization of actions in videos. The state-of-the-art relies on action proposals at test time and selects the best one with a classifier trained on carefully annotated box annotations. Ann…

Action LocalizationMultiple Instance LearningTemporal Localization

Person Re-identification in Videos by Analyzing Spatio-Temporal Tubes

2019-02-13 · Sk. Arif Ahmed, Debi Prosad Dogra, Heeseung Choi, Seungho Chae 외

Typical person re-identification frameworks search for k best matches in a gallery of images that are often collected in varying conditions. The gallery may contain image sequences when re-identification is done on video…

Person Re-IdentificationTemporal Sequences

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation

2024-12-10 · Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T Nguyen 외

To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture tempo…

Contrastive LearningGraph GenerationPanoptic Scene Graph GenerationRelation+2