paper-with-me

Papers

Dynamically Encoded Actions Based on Spacetime Saliency

2015-06-01 · CVPR 2015 6 · Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes

Human actions typically occur over a well localized extent in both space and time. Similarly, as typically captured in video, human actions have small spatiotemporal support in image space. This paper capitalizes on these observations by weighting feature pooling for action recognition over those areas within a video where actions are most likely to occur. To enable this operation, we define a novel measure of spacetime saliency. The measure relies on two observations regarding foreground motion of human actors: They typically exhibit motion that contrasts with that of their surrounding region and they are spatially compact. By using the resulting definition of saliency during feature pooling we show that action recognition performance achieves state-of-the-art levels on three widely considered action recognition datasets. Our saliency weighted pooling can be applied to essentially any locally defined features and encodings thereof. Additionally, we demonstrate that inclusion of locally aggregated spatiotemporal energy features, which efficiently result as a by-product of the saliency computation, further boosts performance over reliance on standard action recognition features alone.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Deep Saliency with Encoded Low level Distance Map and High Level Features

2016-04-19 · CVPR 2016 6 · Gayoung Lee, Yu-Wing Tai, Junmo Kim

Recent advances in saliency detection have utilized deep learning to obtain high level features to detect salient regions in a scene. These advances have demonstrated superior results over previous works that utilize han…

Deep LearningSaliency Detection

Spatio-Temporal Self-Attention Network for Video Saliency Prediction

2021-08-24 · Ziqiang Wang, Zhi Liu, Gongyang Li, Yang Wang 외

3D convolutional neural networks have achieved promising results for video tasks in computer vision, including video saliency prediction that is explored in this paper. However, 3D convolution encodes visual representati…

PredictionSaliency PredictionVideo Saliency Prediction

CaSPR: Learning Canonical Spatiotemporal Point Cloud Representations

2020-08-06 · NeurIPS 2020 12 · Davis Rempe, Tolga Birdal, Yongheng Zhao, Zan Gojcic 외

We propose CaSPR, a method to learn object-centric Canonical Spatiotemporal Point Cloud Representations of dynamically moving or evolving objects. Our goal is to enable information aggregation over time and the interroga…

Camera Pose EstimationObjectPose Estimation

Spatiotemporal Multiplier Networks for Video Action Recognition

2017-07-01 · CVPR 2017 7 · Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes

This paper presents a general ConvNet architecture for video action recognition based on multiplicative interactions of spacetime features. Our model combines the appearance and motion pathways of a two-stream architectu…

Action RecognitionGeneral ClassificationTemporal Action Localization

TSI: Temporal Saliency Integration for Video Action Recognition

2021-06-02 · Haisheng Su, Jinyuan Feng, Dongliang Wang, Weihao Gan 외

Efficient spatiotemporal modeling is an important yet challenging problem for video action recognition. Existing state-of-the-art methods exploit motion clues to assist in short-term temporal modeling through temporal di…

Action RecognitionTemporal Action Localization