Reinforcement Learning-based Mixture of Vision Transformers for Video Violence Recognition
Video violence recognition based on deep learning concerns accurate yet scalable human violence recognition. Currently, most state-of-the-art video violence recognition studies use CNN-based models to represent and categorize videos. However, recent studies suggest that pre-trained transformers are more accurate than CNN-based models on various video analysis benchmarks. Yet these models are not thoroughly evaluated for video violence recognition. This paper introduces a novel transformer-based Mixture of Experts (MoE) video violence recognition system. Through an intelligent combination of large vision transformers and efficient transformer architectures, the proposed system not only takes advantage of the vision transformer architecture but also reduces the cost of utilizing large vision transformers. The proposed architecture maximizes violence recognition system accuracy while actively reducing computational costs through a reinforcement learning-based router. The empirical results show the proposed MoE architecture's superiority over CNN-based models by achieving 92.4% accuracy on the RWF dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Mixture-of-Expertsreinforcement-learningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Video Vision Transformers for Violence Detection
Law enforcement and city safety are significantly impacted by detecting violent incidents in surveillance systems. Although modern (smart) cameras are widely available and affordable, such technological solutions are imp…
Data AugmentationData Efficient Video Transformer for Violence Detection
In smart cities, violence event detection is critical to ensure city safety. Several studies have been done on this topic with a focus on 2d-Convolutional Neural Network (2d-CNN) to detect spatial features from each fram…
Action RecognitionEvent DetectionComparative Analysis: Violence Recognition from Videos using Transfer Learning
Action recognition has become a hot topic in computer vision. However, the main applications of computer vision in video processing have focused on detection of relatively simple actions while complex events such as viol…
Action RecognitionBenchmarkingTransfer LearningVideo Violence Recognition and Localization Using a Semi-Supervised Hard Attention Model
The significant growth of surveillance camera networks necessitates scalable AI solutions to efficiently analyze the large amount of video data produced by these networks. As a typical analysis performed on surveillance …
Activity RecognitionHard Attentionreinforcement-learningReinforcement Learning+1SSIVD-Net: A Novel Salient Super Image Classification & Detection Technique for Weaponized Violence
Detection of violence and weaponized violence in closed-circuit television (CCTV) footage requires a comprehensive approach. In this work, we introduce the \emph{Smart-City CCTV Violence Detection (SCVD)} dataset, specif…
Action Recognitionimage-classificationImage ClassificationVideo Classification+1