paper-with-me

Papers

Tensor Representations for Action Recognition

2020-12-28 · Piotr Koniusz, Lei Wang, Anoop Cherian

Human actions in video sequences are characterized by the complex interplay between spatial features and their temporal dynamics. In this paper, we propose novel tensor representations for compactly capturing such higher-order relationships between visual features for the task of action recognition. We propose two tensor-based feature representations, viz. (i) sequence compatibility kernel (SCK) and (ii) dynamics compatibility kernel (DCK). SCK builds on the spatio-temporal correlations between features, whereas DCK explicitly models the action dynamics of a sequence. We also explore generalization of SCK, coined SCK(+), that operates on subsequences to capture the local-global interplay of correlations, which can incorporate multi-modal inputs e.g., skeleton 3D body-joints and per-frame classifier scores obtained from deep learning models trained on videos. We introduce linearization of these kernels that lead to compact and fast descriptors. We provide experiments on (i) 3D skeleton action sequences, (ii) fine-grained video sequences, and (iii) standard non-fine-grained videos. As our final representations are tensors that capture higher-order relationships of features, they relate to co-occurrences for robust fine-grained recognition. We use higher-order tensors and so-called Eigenvalue Power Normalization (EPN) which have been long speculated to perform spectral detection of higher-order occurrences, thus detecting fine-grained relationships of features rather than merely count features in action sequences. We prove that a tensor of order r, built from Z* dimensional features, coupled with EPN indeed detects if at least one higher-order occurrence is projected' into one of its binom(Z*,r) subspaces of dim. r represented by the tensor, thus forming a Tensor Power Normalization metric endowed with binom(Z*,r) such detectors'.

📄 PDF Abstract BibTeX arXiv:2012.14371

Code (1)

ElricXin/Skeleton-MixFormer pytorch

Tasks

Action RecognitionAction Recognition In VideosSkeleton Based Action Recognition

Similar Papers 제목 키워드 기반

Tensor Representations via Kernel Linearization for Action Recognition from 3D Skeletons (Extended Version)

2016-04-01 · Piotr Koniusz, Anoop Cherian, Fatih Porikli

In this paper, we explore tensor representations that can compactly capture higher-order relationships between skeleton joints for 3D action recognition. We first define RBF kernels on 3D joint sequences, which are then …

3D Action RecognitionAction RecognitionFormTemporal Action Localization

Topic Tensor Network for Implicit Discourse Relation Recognition in Chinese

2019-07-01 · ACL 2019 7 · Sheng Xu, Peifeng Li, Fang Kong, Qiaoming Zhu 외

In the literature, most of the previous studies on English implicit discourse relation recognition only use sentence-level representations, which cannot provide enough semantic information in Chinese due to its unique pa…

RelationSentenceTensor Networks

Implicit Discourse Relation Recognition using Neural Tensor Network with Interactive Attention and Sparse Learning

2018-08-01 · COLING 2018 8 · Fengyu Guo, Ruifang He, Di Jin, Jianwu Dang 외

Implicit discourse relation recognition aims to understand and annotate the latent relations between two discourse arguments, such as temporal, comparison, etc. Most previous methods encode two discourse arguments separa…

RelationSparse LearningText Summarization

Large Margin Low Rank Tensor Analysis

2013-06-11 · Guoqiang Zhong, Mohamed Cheriet

Other than vector representations, the direct objects of human cognition are generally high-order tensors, such as 2D images and 3D textures. From this fact, two interesting questions naturally arise: How does the human …

Face RecognitionObject Recognition

Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization

2019-07-01 · ACL 2019 7 · Paul Pu Liang, Zhun Liu, Yao-Hung Hubert Tsai, Qibin Zhao 외

There has been an increased interest in multimodal language processing including multimodal dialog, question answering, sentiment analysis, and speech recognition. However, naturally occurring multimodal data is often im…

Question AnsweringSentiment Analysisspeech-recognitionSpeech Recognition+2