GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
Video anomaly detection (VAD) is a challenging task that detects anomalous frames in continuous surveillance videos. Most previous work utilizes the spatio-temporal correlation of visual features to distinguish whether there are abnormalities in video snippets. Recently, some works attempt to introduce multi-modal information, like text feature, to enhance the results of video anomaly detection. However, these works merely incorporate text features into video snippets in a coarse manner, overlooking the significant amount of redundant information that may exist within the video snippets. Therefore, we propose to leverage the diversity among multi-modal information to further refine the extracted features, reducing the redundancy in visual features, and we propose Grained Multi-modal Feature for Video Anomaly Detection (GMFVAD). Specifically, we generate more grained multi-modal feature based on the video snippet, which summarizes the main content, and text features based on the captions of original video will be introduced to further enhance the visual features of highlighted portions. Experiments show that the proposed GMFVAD achieves state-of-the-art performance on four mainly datasets. Ablation experiments also validate that the improvement of GMFVAD is due to the reduction of redundant information.
Code (0)
등록된 구현이 없습니다.
Tasks
Video Anomaly DetectionSimilar Papers 제목 키워드 기반
NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-ID
Multi-modal object re-identification (ReID) aims to extract identity features across heterogeneous spectral modalities to enable accurate recognition and retrieval in complex real-world scenarios. However, most existing …
AttributeCaption GenerationDescriptiveMixture-of-Experts+1Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
Recent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-modal local representations, namely frame…
cross-modal alignmentFeature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detection
Multimodal methods are widely used in rice deterioration detection, but they exhibit limited capability in representing and extracting fine-grained abnormal features. Moreover, these methods rely on devices such as hyper…
Multi-modal Fake News Detection on Social Media via Multi-grained Information Fusion
The easy sharing of multimedia content on social media has caused a rapid dissemination of fake news, which threatens society's stability and security. Therefore, fake news detection has garnered extensive research inter…
Fake News DetectionMCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification
Multimodal information processing has become increasingly important for enhancing image classification performance. However, the intricate and implicit dependencies across different modalities often hinder conventional m…
Classificationimage-classificationImage Classification