Localized Trajectories for 2D and 3D Action Recognition
The Dense Trajectories concept is one of the most successful approaches in action recognition, suitable for scenarios involving a significant amount of motion. However, due to noise and background motion, many generated trajectories are irrelevant to the actual human activity and can potentially lead to performance degradation. In this paper, we propose Localized Trajectories as an improved version of Dense Trajectories where motion trajectories are clustered around human body joints provided by RGB-D cameras and then encoded by local Bag-of-Words. As a result, the Localized Trajectories concept provides a more discriminative representation of actions as compared to Dense Trajectories. Moreover, we generalize Localized Trajectories to 3D by using the modalities offered by RGB-D cameras. One of the main advantages of using RGB-D data to generate trajectories is that they include radial displacements that are perpendicular to the image plane. Extensive experiments and analysis are carried out on five different datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Action RecognitionAction RecognitionTemporal Action LocalizationSimilar Papers 제목 키워드 기반
AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions
This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in sp…
Actin DetectionAction DetectionAction LocalizationAction Recognition+3TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text co…
Text-to-Video GenerationHandcrafted localized phase features for human action recognition
Human action recognition is one of the most important topics in computer vision. Monitoring elderly people and children, smart surveillance systems and human-computer interaction are a few examples of its applications. T…
Action ClassificationAction RecognitionOptical Flow EstimationTemporal Action LocalizationTemporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images
We address the problem of fine-grained action localization from temporally untrimmed web videos. We assume that only weak video-level annotations are available for training. The goal is to use these weak labels to identi…
Action LocalizationAction RecognitionTemporal Action LocalizationTemporal LocalizationSpan-based Joint Entity and Relation Extraction with Transformer Pre-training
We introduce SpERT, an attention model for span-based joint entity and relation extraction. Our key contribution is a light-weight reasoning on BERT embeddings, which features entity recognition and filtering, as well as…
Joint Entity and Relation ExtractionNamed Entity Recognition (NER)RelationRelation Classification+2