RubiksNet: Learnable 3D-Shift for Efficient Video Action Recognition
Video action recognition is a complex task dependent on modeling spatial and temporal context. Standard approaches rely on 2D or 3D convolutions to process such context, resulting in expensive operations with millions of parameters. Recent efficient architectures leverage a channel-wise shift-based primitive as a replacement for temporal convolutions, but remain bottlenecked by spatial convolution operations to maintain strong accuracy and a fixed-shift scheme. Naively extending such developments to a 3D setting is a difficult, intractable goal. To this end, we introduce RubiksNet, a new efficient architecture for video action recognition which is based on a proposed learnable 3D spatiotemporal shift operation instead. We analyze the suitability of our new primitive for video action recognition and explore several novel variations of our approach to enable stronger representational flexibility while maintaining an efficient design. We benchmark our approach on several standard video recognition datasets, and observe that our method achieves comparable or better accuracy than prior work on efficient video action recognition at a fraction of the performance cost, with 2.9-5.9x fewer parameters and 2.1-3.7x fewer FLOPs. We also perform a series of controlled ablation studies to verify our significant boost in the efficiency-accuracy tradeoff curve is rooted in the core contributions of our RubiksNet architecture.
Code (1)
Tasks
Action RecognitionTemporal Action LocalizationVideo RecognitionSimilar Papers 제목 키워드 기반
Learnable Sampling 3D Convolution for Video Enhancement and Action Recognition
A key challenge in video enhancement and action recognition is to fuse useful information from neighboring frames. Recent works suggest establishing accurate correspondences between neighboring frames before fusing tempo…
Action RecognitionDenoisingSuper-ResolutionVideo Denoising+2Learnable Motion-Focused Tokenization for Effective and Efficient Video Unsupervised Domain Adaptation
Video Unsupervised Domain Adaptation (VUDA) poses a significant challenge in action recognition, requiring the adaptation of a model from a labeled source domain to an unlabeled target domain. Despite recent advances, ex…
Unsupervised Domain AdaptationComputational EfficiencyAction RecognitionOnline Learnable Keyframe Extraction in Videos and its Application with Semantic Word Vector in Action Recognition
Video processing has become a popular research direction in computer vision due to its various applications such as video summarization, action recognition, etc. Recently, deep learning-based methods have achieved impres…
Action RecognitionGeneral ClassificationVideo SummarizationDVANet: Disentangling View and Action Features for Multi-View Action Recognition
In this work, we present a novel approach to multi-view action recognition where we guide learned action representations to be separated from view-relevant information in a video. When trying to classify action instances…
Action RecognitionAction Recognition In VideosDecoderHigher-order Network for Action Recognition
Capturing spatiotemporal dynamics is an essential topic in video recognition. In this paper, we present learnable higher-order operations as a generic family of building blocks for capturing spatiotemporal dynamics from …
Action RecognitionGeneral ClassificationVideo ClassificationVideo Recognition