Audio-Guided Attention Network for Weakly Supervised Violence Detection
Detecting violence in video is a challenging task due to its complex scenarios and great intra-class variability. Most previous works specialize in the analysis of appearance or motion information, ignoring the co-occurrence of some audio and visual events. Physical conflicts such as abuse and fighting are usually accompanied by screaming, while crowd violence such as riots and wars are generally related to gunshots and explosions. Therefore, we propose a novel audio-guided multimodal violence detection framework. First, deep neural networks are used to extract appearance and audio features, respectively. Then, a Cross-Modal Awareness Local-Arousal (CMA-LA) network is proposed for cross-modal interaction, which implements audio-to-visual feature enhancement over temporal dimension. The enhanced features are then fed into a multilayer perceptron (MLP) to capture high-level semantics, followed by a temporal convolution layer to obtain high-confidence violence scores. To validate the proposed method, we conduct experiments on a large violent video dataset, XD Violence. Comprehensive experiments demonstrate the robust performance of our approach, which also achieves a new state-of-the-art AP result.
Code (1)
Tasks
Anomaly Detection In Surveillance VideosMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Modality-Aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection
Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early …
Anomaly Detection In Surveillance Videosaudio-visual learningMultiple Instance LearningLearning Weakly Supervised Audio-Visual Violence Detection in Hyperbolic Space
In recent years, the task of weakly supervised audio-visual violence detection has gained considerable attention. The goal of this task is to identify violent segments within multimodal data based on video-level labels. …
Anomaly Detection In Surveillance VideosCross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
Recently, weakly supervised video anomaly detection (WS-VAD) has emerged as a contemporary research direction to identify anomaly events like violence and nudity in videos using only video-level labels. However, this tas…
Anomaly DetectionGraph AttentionVideo Anomaly DetectionWeakly-supervised Video Anomaly DetectionMulti-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
Weakly supervised multimodal violence detection aims to learn a violence detection model by leveraging multiple modalities such as RGB, optical flow, and audio, while only video-level annotations are available. In the pu…
Anomaly Detection In Surveillance VideosOptical Flow EstimationAligning First, Then Fusing: A Novel Weakly Supervised Multimodal Violence Detection Method
Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrate…
Anomaly Detection In Surveillance VideosMultiple Instance LearningOptical Flow Estimation