Papers Spatio-Temporal Action Localization
“Spatio-Temporal Action Localization” 태그가 달린 논문 39편 · 필터 해제
Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining
We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target …
Spatio-Temporal Action LocalizationContrastive LearningLearning from Synthetic Data via Provenance-Based Input Gradient Guidance
Learning methods using synthetic data have attracted attention as an effective approach for increasing the diversity of training data while reducing collection costs, thereby improving the robustness of model discriminat…
Spatio-Temporal Action LocalizationImage ClassificationObject LocalizationScaling Open-Vocabulary Action Detection
In this work, we focus on scaling open-vocabulary action detection. Existing approaches for action detection are predominantly limited to closed-set scenarios and rely on complex, parameter-heavy architectures. Extending…
Action DetectionMultiple Action DetectionOpen Vocabulary Action DetectionSpatio-Temporal Action Localization+2Minimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues
This paper introduces ViNet-S, a 36MB model based on the ViNet architecture with a U-Net design, featuring a lightweight decoder that significantly reduces model size and parameters without compromising performance. Addi…
Action ClassificationAction LocalizationDecoderSaliency Prediction+3Survey of Action Recognition, Spotting and Spatio-Temporal Localization in Soccer -- Current Trends and Research Perspectives
Action scene understanding in soccer is a challenging task due to the complex and dynamic nature of the game, as well as the interactions between players. This article provides a comprehensive overview of this task divid…
Action LocalizationAction RecognitionScene UnderstandingSpatio-Temporal Action Localization+2End-to-End Spatio-Temporal Action Localisation with Video Transformers
The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end, purely-transformer based model that directly ingests an input…
Action DetectionAction RecognitionSpatio-Temporal Action LocalizationVideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
Scale is the primary factor for building a powerful foundation model that could well generalize to a variety of downstream tasks. However, it is still challenging to train video foundation models with billions of paramet…
Action ClassificationAction RecognitionAction Recognition In VideosDecoder+3Unmasked Teacher: Towards Training-Efficient Video Foundation Models
Video Foundation Models (VFMs) have received limited exploration due to high computational costs and data scarcity. Previous VFMs rely on Image Foundation Models (IFMs), which face challenges in transferring to the video…
Action ClassificationAction Recognitioncross-modal alignmentSpatio-Temporal Action Localization+4Unified Keypoint-based Action Recognition Framework via Structured Keypoint Pooling
This paper simultaneously addresses three limitations associated with conventional skeleton-based action recognition; skeleton detection and tracking errors, poor variety of the targeted actions, as well as person-wise a…
Action LocalizationAction RecognitionActivity RecognitionData Augmentation+6InternVideo: General Video Foundation Models via Generative and Discriminative Learning
The foundation models have recently shown excellent performance on a variety of downstream tasks in computer vision. However, most existing vision foundation models simply focus on image-level pretraining and adpation, w…
Action ClassificationAction RecognitionContrastive LearningOpen Set Action Recognition+8E^2TAD: An Energy-Efficient Tracking-based Action Detector
Video action detection (spatio-temporal action localization) is usually the starting point for human-centric intelligent analysis of videos nowadays. It has high practical impacts for many applications across robotics, s…
Action DetectionAction LocalizationFine-Grained Action Detectionobject-detection+4MM-SEAL: A Large-scale Video Dataset of Multi-person Multi-grained Spatio-temporally Action Localization
In this paper, we introduce a novel large-scale video dataset dubbed MM-SEAL for multi-person multi-grained spatio-temporal action localization among human daily life. We are the first to propose a new benchmark for mult…
Action LocalizationAction RecognitionSpatio-Temporal Action LocalizationTemporal Action Localization+1Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision
Modern self-supervised learning algorithms typically enforce persistency of instance representations across views. While being very effective on learning holistic image and video representations, such an objective become…
Action LocalizationAction RecognitionContrastive LearningObject Tracking+3KORSAL: Key-point Detection based Online Real-Time Spatio-Temporal Action Localization
Real-time and online action localization in a video is a critical yet highly challenging problem. Accurate action localization requires the utilization of both temporal and spatial information. Recent attempts achieve th…
Action LocalizationOptical Flow EstimationSpatio-Temporal Action LocalizationTemporal Action LocalizationRelation Modeling in Spatio-Temporal Action Localization
This paper presents our solution to the AVA-Kinetics Crossover Challenge of ActivityNet workshop at CVPR 2021. Our solution utilizes multiple types of relation modeling methods for spatio-temporal action detection and ad…
Action DetectionAction LocalizationRelationSpatio-Temporal Action Localization+1ST-HOI: A Spatial-Temporal Baseline for Human-Object Interaction Detection in Videos
Detecting human-object interactions (HOI) is an important step toward a comprehensive visual understanding of machines. While detecting non-temporal HOIs (e.g., sitting on a chair) from static images is feasible, it is u…
Action DetectionHuman-Object Interaction AnticipationHuman-Object Interaction DetectionSpatio-Temporal Action LocalizationRelevance Detection in Cataract Surgery Videos by Spatio-Temporal Action Localization
In cataract surgery, the operation is performed with the help of a microscope. Since the microscope enables watching real-time surgery by up to two people only, a major part of surgical training is conducted using the re…
Action LocalizationRelevance DetectionRetrievalSpatio-Temporal Action Localization+1Real-time Spatio-temporal Action Localization via Learning Motion Representation
Abstract. Most state-of-the-art spatio-temporal (S-T) action localization methods explicitly use optical flow as auxiliary motion information. Although the combination of optical flow and RGB significantly improves the p…
Action ClassificationAction LocalizationKnowledge DistillationOptical Flow Estimation+2Unsupervised Domain Adaptation for Spatio-Temporal Action Localization
Spatio-temporal action localization is an important problem in computer vision that involves detecting where and when activities occur, and therefore requires modeling of both spatial and temporal features. This problem …
Action LocalizationDomain Adaptationobject-detectionObject Detection+3CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Most current pipelines for spatio-temporal action localization connect frame-wise or clip-wise detection results to generate action proposals, where only local information is exploited and the efficiency is hindered by d…
Action DetectionAction LocalizationSpatio-Temporal Action LocalizationTemporal Action Localization