AE-Net:Adjoint Enhancement Network for Efficient Action Recognition in Video Understanding
Action recognition in video understanding is a challenging task, largely because of the complexity and difficulty in temporal modeling, making it suffer from motion information loss and misalignment of temporal attention in spatial dimensions. To overcome these difficulties, we propose a novel temporal modeling method called Adjoint Enhancement Network (AE-Net), which can fully explore clues of motion and time in the long-range structure. The AE-Net mainly consists of two new modules: the Initial Adjoint Enhancement Module (IAE-Module), which deals with shallow features; and the Global Adjoint Enhancement Module (GAE-Module), which deals with global features. With a novel mechanism of parallel spatio-temporal convolution and difference fusion, the IAE-Module is to enhance the degree of motion transformation in shallow network features, exciting the potential of motion flow and avoiding motion information loss. The GAE-Module is proposed to improve the local temporal representation in long-range structures by feeding the enhanced feature differences into a spatial cascade module with residuals to resolve the misalignment of temporal attention in the spatial dimension.The experimental results show that our AE-Net can achieve state-of-the-art results in Something-Something V1, UCF101 and HMDB-51 datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionVideo UnderstandingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ARID: A New Dataset for Recognizing Action in the Dark
The task of action recognition in dark videos is useful in various scenarios, e.g., night surveillance and self-driving at night. Though progress has been made in the action recognition task for videos in normal illumina…
Action RecognitionCross-Enhancement Transform Two-Stream 3D ConvNets for Action Recognition
Action recognition is an important research topic in computer vision. It is the basic work for visual understanding and has been applied in many fields. Since human actions can vary in different environments, it is diffi…
Action RecognitionAutonomous DrivingAutonomous VehiclesOptical Flow Estimation+2The Role of Video Generation in Enhancing Data-Limited Action Understanding
Video action understanding tasks in real-world scenarios always suffer data limitations. In this paper, we address the data-limited action understanding problem by bridging data scarcity. We propose a novel method that e…
Action RecognitionAction UnderstandingVideo GenerationZero-Shot Action RecognitionIndGIC: Supervised Action Recognition under Low Illumination
Technologies of human action recognition in the dark are gaining more and more attention as huge demand in surveillance, motion control and human-computer interaction. However, because of limitation in image enhancement …
Action RecognitionImage EnhancementTemporal Action LocalizationLearnable Sampling 3D Convolution for Video Enhancement and Action Recognition
A key challenge in video enhancement and action recognition is to fuse useful information from neighboring frames. Recent works suggest establishing accurate correspondences between neighboring frames before fusing tempo…
Action RecognitionDenoisingSuper-ResolutionVideo Denoising+2