paper-with-me

홈 › Papers

GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection

2025-10-23 · Guangyu Dai, Dong Chen, Siliang Tang, Yueting Zhuang arxiv

Video anomaly detection (VAD) is a challenging task that detects anomalous frames in continuous surveillance videos. Most previous work utilizes the spatio-temporal correlation of visual features to distinguish whether there are abnormalities in video snippets. Recently, some works attempt to introduce multi-modal information, like text feature, to enhance the results of video anomaly detection. However, these works merely incorporate text features into video snippets in a coarse manner, overlooking the significant amount of redundant information that may exist within the video snippets. Therefore, we propose to leverage the diversity among multi-modal information to further refine the extracted features, reducing the redundancy in visual features, and we propose Grained Multi-modal Feature for Video Anomaly Detection (GMFVAD). Specifically, we generate more grained multi-modal feature based on the video snippet, which summarizes the main content, and text features based on the captions of original video will be introduced to further enhance the visual features of highlighted portions. Experiments show that the proposed GMFVAD achieves state-of-the-art performance on four mainly datasets. Ablation experiments also validate that the improvement of GMFVAD is due to the reduction of redundant information.

📄 PDF Abstract BibTeX arXiv:2510.20268

Code (0)

등록된 구현이 없습니다.

Tasks

Video Anomaly Detection

Similar Papers 제목 키워드 기반

NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-ID

2025-05-26 · Shihao Li, Chenglong Li, Aihua Zheng, Andong Lu 외

Multi-modal object re-identification (ReID) aims to extract identity features across heterogeneous spectral modalities to enable accurate recognition and retrieval in complex real-world scenarios. However, most existing …

AttributeCaption GenerationDescriptiveMixture-of-Experts+1

Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training

2024-08-15 · Yiming Li, Zhifang Guo, Xiangdong Wang, Hong Liu

Recent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-modal local representations, namely frame…

cross-modal alignment

Feature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detection

2026-02-16 · Rongqiang Zhao, Hengrui Hu, Yijing Wang, Mingchun Sun 외 arxiv

Multimodal methods are widely used in rice deterioration detection, but they exhibit limited capability in representing and extracting fine-grained abnormal features. Moreover, these methods rely on devices such as hyper…

Multi-modal Fake News Detection on Social Media via Multi-grained Information Fusion

2023-04-03 · Yangming Zhou, Yuzhou Yang, Qichao Ying, Zhenxing Qian 외

The easy sharing of multimedia content on social media has caused a rapid dissemination of fake news, which threatens society's stability and security. Therefore, fake news detection has garnered extensive research inter…

Fake News Detection

MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification

2025-05-29 · Yang Qiao, Xiaoyu Zhong, Xiaofeng Gu, Zhiguo Yu

Multimodal information processing has become increasingly important for enhancing image classification performance. However, the intricate and implicit dependencies across different modalities often hinder conventional m…

Classificationimage-classificationImage Classification