Learning Weakly Supervised Audio-Visual Violence Detection in Hyperbolic Space
In recent years, the task of weakly supervised audio-visual violence detection has gained considerable attention. The goal of this task is to identify violent segments within multimodal data based on video-level labels. Despite advances in this field, traditional Euclidean neural networks, which have been used in prior research, encounter difficulties in capturing highly discriminative representations due to limitations of the feature space. To overcome this, we propose HyperVD, a novel framework that learns snippet embeddings in hyperbolic space to improve model discrimination. Our framework comprises a detour fusion module for multimodal fusion, effectively alleviating modality inconsistency between audio and visual signals. Additionally, we contribute two branches of fully hyperbolic graph convolutional networks that excavate feature similarities and temporal relationships among snippets in hyperbolic space. By learning snippet representations in this space, the framework effectively learns semantic discrepancies between violent and normal events. Extensive experiments on the XD-Violence benchmark demonstrate that our method outperforms state-of-the-art methods by a sizable margin.
Code (1)
Tasks
Anomaly Detection In Surveillance VideosSimilar Papers 제목 키워드 기반
Modality-Aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection
Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early …
Anomaly Detection In Surveillance Videosaudio-visual learningMultiple Instance LearningAudio-Guided Attention Network for Weakly Supervised Violence Detection
Detecting violence in video is a challenging task due to its complex scenarios and great intra-class variability. Most previous works specialize in the analysis of appearance or motion information, ignoring the co-occurr…
Anomaly Detection In Surveillance VideosMulti-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
Weakly supervised multimodal violence detection aims to learn a violence detection model by leveraging multiple modalities such as RGB, optical flow, and audio, while only video-level annotations are available. In the pu…
Anomaly Detection In Surveillance VideosOptical Flow EstimationAligning First, Then Fusing: A Novel Weakly Supervised Multimodal Violence Detection Method
Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrate…
Anomaly Detection In Surveillance VideosMultiple Instance LearningOptical Flow EstimationCross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
Recently, weakly supervised video anomaly detection (WS-VAD) has emerged as a contemporary research direction to identify anomaly events like violence and nudity in videos using only video-level labels. However, this tas…
Anomaly DetectionGraph AttentionVideo Anomaly DetectionWeakly-supervised Video Anomaly Detection