paper-with-me

홈 › Papers

Second-order Temporal Pooling for Action Recognition

2017-04-23 · Anoop Cherian, Stephen Gould

Deep learning models for video-based action recognition usually generate features for short clips (consisting of a few frames); such clip-level features are aggregated to video-level representations by computing statistics on these features. Typically zero-th (max) or the first-order (average) statistics are used. In this paper, we explore the benefits of using second-order statistics. Specifically, we propose a novel end-to-end learnable feature aggregation scheme, dubbed temporal correlation pooling that generates an action descriptor for a video sequence by capturing the similarities between the temporal evolution of clip-level CNN features computed across the video. Such a descriptor, while being computationally cheap, also naturally encodes the co-activations of multiple CNN features, thereby providing a richer characterization of actions than their first-order counterparts. We also propose higher-order extensions of this scheme by computing correlations after embedding the CNN features in a reproducing kernel Hilbert space. We provide experiments on benchmark datasets such as HMDB-51 and UCF-101, fine-grained datasets such as MPII Cooking activities and JHMDB, as well as the recent Kinetics-600. Our results demonstrate the advantages of higher-order pooling schemes that when combined with hand-crafted features (as is standard practice) achieves state-of-the-art accuracy.

📄 PDF Abstract BibTeX arXiv:1704.06925

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

3Mformer: Multi-order Multi-mode Transformer for Skeletal Action Recognition

2023-03-25 · CVPR 2023 1 · Lei Wang, Piotr Koniusz

Many skeletal action recognition models use GCNs to represent the human body by 3D body joints connected body parts. GCNs aggregate one- or few-hop graph neighbourhoods, and ignore the dependency between not linked body …

Action RecognitionSkeleton Based Action Recognition

Order-aware Convolutional Pooling for Video Based Action Recognition

2016-01-31 · Peng Wang, Lingqiao Liu, Chunhua Shen, Heng Tao Shen

Most video based action recognition approaches create the video-level representation by temporally pooling the features extracted at each frame. The pooling methods that they adopt, however, usually completely or partial…

Action RecognitionTemporal Action Localization

Neural Koopman Pooling: Control-Inspired Temporal Dynamics Encoding for Skeleton-Based Action Recognition

2023-01-01 · CVPR 2023 1 · Xinghan Wang, Xin Xu, Yadong Mu

Skeleton-based human action recognition is becoming increasingly important in a variety of fields. Most existing works train a CNN or GCN based backbone to extract spatial-temporal features, and use temporal average/…

Action RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Rank Pooling for Action Recognition

2015-12-06 · Basura Fernando, Efstratios Gavves, Jose Oramas, Amir Ghodrati 외

We propose a function-based temporal pooling method that captures the latent structure of the video sequence data - e.g. how frame-level features evolve over time in a video. We show how the parameters of a function that…

Action RecognitionGesture RecognitionLearning-To-RankTemporal Action Localization

Voronoi-based Second-order Descriptor with Whitened Metric in LiDAR Place Recognition

2026-03-16 · Jaein Kim, Hee Bin Yoo, Dong-Sig Han, Byoung-Tak Zhang arxiv

The pooling layer plays a vital role in aggregating local descriptors into the metrizable global descriptor in the LiDAR Place Recognition (LPR). In particular, the second-order pooling is capable of capturing higher-ord…