paper-with-me

홈 › Papers

Audio-Guided Attention Network for Weakly Supervised Violence Detection

2022-02-21 · Conference 2022 2 · Yujiang Pu, Xiaoyu Wu

Detecting violence in video is a challenging task due to its complex scenarios and great intra-class variability. Most previous works specialize in the analysis of appearance or motion information, ignoring the co-occurrence of some audio and visual events. Physical conflicts such as abuse and fighting are usually accompanied by screaming, while crowd violence such as riots and wars are generally related to gunshots and explosions. Therefore, we propose a novel audio-guided multimodal violence detection framework. First, deep neural networks are used to extract appearance and audio features, respectively. Then, a Cross-Modal Awareness Local-Arousal (CMA-LA) network is proposed for cross-modal interaction, which implements audio-to-visual feature enhancement over temporal dimension. The enhanced features are then fed into a multilayer perceptron (MLP) to capture high-level semantics, followed by a temporal convolution layer to obtain high-confidence violence scores. To validate the proposed method, we conduct experiments on a large violent video dataset, XD Violence. Comprehensive experiments demonstrate the robust performance of our approach, which also achieves a new state-of-the-art AP result.

📄 PDF Abstract BibTeX

Code (1)

Aaron-Pu/cma_xdVioDet 공식 구현 pytorch

Tasks

Anomaly Detection In Surveillance Videos

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Modality-Aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection

2022-07-12 · Jiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng 외

Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early …

Anomaly Detection In Surveillance Videosaudio-visual learningMultiple Instance Learning

Learning Weakly Supervised Audio-Visual Violence Detection in Hyperbolic Space

2023-05-30 · Xiaogang Peng, Hao Wen, Yikai Luo, Xiao Zhou 외

In recent years, the task of weakly supervised audio-visual violence detection has gained considerable attention. The goal of this task is to identify violent segments within multimodal data based on video-level labels. …

Anomaly Detection In Surveillance Videos

Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection

2024-12-29 · CVPR 2024 6 · Ayush Ghadiya, Purbayan Kar, Vishal Chudasama, Pankaj Wasnik

Recently, weakly supervised video anomaly detection (WS-VAD) has emerged as a contemporary research direction to identify anomaly events like violence and nudity in videos using only video-level labels. However, this tas…

Anomaly DetectionGraph AttentionVideo Anomaly DetectionWeakly-supervised Video Anomaly Detection

Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection

2024-05-08 · Shengyang Sun, Xiaojin Gong

Weakly supervised multimodal violence detection aims to learn a violence detection model by leveraging multiple modalities such as RGB, optical flow, and audio, while only video-level annotations are available. In the pu…

Anomaly Detection In Surveillance VideosOptical Flow Estimation

Aligning First, Then Fusing: A Novel Weakly Supervised Multimodal Violence Detection Method

2025-01-13 · Wenping Jin, Li Zhu, Jing Sun

Weakly supervised violence detection refers to the technique of training models to identify violent segments in videos using only video-level labels. Among these approaches, multimodal violence detection, which integrate…

Anomaly Detection In Surveillance VideosMultiple Instance LearningOptical Flow Estimation