AM Flow: Adapters for Temporal Processing in Action Recognition
Deep learning models, in particular \textit{image} models, have recently gained generalisability and robustness. %are becoming more general and robust by the day. In this work, we propose to exploit such advances in the realm of \textit{video} classification. Video foundation models suffer from the requirement of extensive pretraining and a large training time. Towards mitigating such limitations, we propose "\textit{Attention Map (AM) Flow}" for image models, a method for identifying pixels relevant to motion in each input video frame. In this context, we propose two methods to compute AM flow, depending on camera motion. AM flow allows the separation of spatial and temporal processing, while providing improved results over combined spatio-temporal processing (as in video models). Adapters, one of the popular techniques in parameter efficient transfer learning, facilitate the incorporation of AM flow into pretrained image models, mitigating the need for full-finetuning. We extend adapters to "\textit{temporal processing adapters}" by incorporating a temporal processing unit into the adapters. Our work achieves faster convergence, therefore reducing the number of epochs needed for training. Moreover, we endow an image model with the ability to achieve state-of-the-art results on popular action recognition datasets. This reduces training time and simplifies pretraining. We present experiments on Kinetics-400, Something-Something v2, and Toyota Smarthome datasets, showcasing state-of-the-art or comparable results.
Code (0)
등록된 구현이 없습니다.
Tasks
Action ClassificationAction RecognitionTransfer LearningVideo ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Frame2Freq: Spectral Adapters for Fine-Grained Video Understanding
Adapting image-pretrained backbones to video typically relies on time-domain adapters tuned to a single temporal scale. Our experiments show that these modules pick up static image cues and very fast flicker changes, whi…
Activity RecognitionAction RecognitionELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
Self-supervised learning has emerged as a key approach for learning generic representations from speech data. Despite promising results in downstream tasks such as speech recognition, speaker verification, and emotion re…
Emotion Recognitionparameter-efficient fine-tuningSelf-Supervised LearningSpeaker Verification+2LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
Temporal Action Localization (TAL) involves localizing and classifying action snippets in an untrimmed video. The emergence of large video foundation models has led RGB-only video backbones to outperform previous methods…
Action LocalizationGPUOptical Flow EstimationTemporal Action Localization+1Decoupled Prompt-Adapter Tuning for Continual Activity Recognition
Action recognition technology plays a vital role in enhancing security through surveillance systems, enabling better patient monitoring in healthcare, providing in-depth performance analysis in sports, and facilitating s…
Action RecognitionActivity RecognitionCombining Spatio-Temporal Appearance Descriptors and Optical Flow for Human Action Recognition in Video Data
This paper proposes combining spatio-temporal appearance (STA) descriptors with optical flow for human action recognition. The STA descriptors are local histogram-based descriptors of space-time, suitable for building a …
Action RecognitionOptical Flow EstimationTemporal Action Localization