Temporal Difference Networks for Action Recognition
Temporal modeling still remains challenging for action recognition in videos. To mitigate this issue, this paper presents a new video architecture, termed as Temporal Difference Network (TDN), with a focus on capturing multi-scale temporal information for efficient action recognition. The core of our TDN is to devise an efficient temporal module (TDM) by explicitly leveraging a temporal difference operator, and systematically assess its effect on short-term and long-term motion modeling. To fully capture temporal information over the entire video, our TDN is established with a two-level difference modeling paradigm. Specifically, for local motion modeling, temporal difference over consecutive frames is used to supply 2D CNNs with finer motion pattern, while for global motion modeling, temporal difference across segments is incorporated to capture long-range structure for motion feature excitation. TDN provides a simple and principled temporal modeling framework, and could be instantiated with the existing CNNs at a small extra computational cost. Our TDN presents a new state of the art on the datasets of Something-Something V1 \& V2 and Kinetics-400 under the setting of using similar backbones. In addition, we present some visualization results on our TDN and try to provide new insights on temporal difference operation.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction Recognition In VideosMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
TDN: Temporal Difference Networks for Efficient Action Recognition
Temporal modeling still remains challenging for action recognition in videos. To mitigate this issue, this paper presents a new video architecture, termed as Temporal Difference Network (TDN), with a focus on capturing m…
Action ClassificationAction RecognitionAction Recognition In VideosGlobal Temporal Difference Network for Action Recognition
—Temporal modeling still remains as a challenge for action recognition. Most existing temporal models focus on learning local variation between neighbor frames. There exists obvious deviations between local and global…
Action RecognitionDynamic Spatio-Temporal Specialization Learning for Fine-Grained Action Recognition
The goal of fine-grained action recognition is to successfully discriminate between action categories with subtle differences. To tackle this, we derive inspiration from the human visual system which contains specialized…
Action RecognitionFine-grained Action RecognitionTDS-CLIP: Temporal Difference Side Network for Image-to-Video Transfer Learning
Recently, large-scale pre-trained vision-language models (e.g., CLIP), have garnered significant attention thanks to their powerful representative capabilities. This inspires researchers in transferring the knowledge fro…
Action Recognitionparameter-efficient fine-tuningTemporal Action LocalizationTransfer LearningMulti-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture Recognition
RGB-D action and gesture recognition remain an interesting topic in human-centered scene understanding, primarily due to the multiple granularities and large variation in human motion. Although many RGB-D based action an…
Gesture RecognitionScene Understanding