paper-with-me

Papers

AssistPDA: An Online Video Surveillance Assistant for Video Anomaly Prediction, Detection, and Analysis

2025-03-27 · Zhiwei Yang, Chen Gao, Jing Liu, Peng Wu, Guansong Pang, Mike Zheng Shou

The rapid advancements in large language models (LLMs) have spurred growing interest in LLM-based video anomaly detection (VAD). However, existing approaches predominantly focus on video-level anomaly question answering or offline detection, ignoring the real-time nature essential for practical VAD applications. To bridge this gap and facilitate the practical deployment of LLM-based VAD, we introduce AssistPDA, the first online video anomaly surveillance assistant that unifies video anomaly prediction, detection, and analysis (VAPDA) within a single framework. AssistPDA enables real-time inference on streaming videos while supporting interactive user engagement. Notably, we introduce a novel event-level anomaly prediction task, enabling proactive anomaly forecasting before anomalies fully unfold. To enhance the ability to model intricate spatiotemporal relationships in anomaly events, we propose a Spatio-Temporal Relation Distillation (STRD) module. STRD transfers the long-term spatiotemporal modeling capabilities of vision-language models (VLMs) from offline settings to real-time scenarios. Thus it equips AssistPDA with a robust understanding of complex temporal dependencies and long-sequence memory. Additionally, we construct VAPDA-127K, the first large-scale benchmark designed for VLM-based online VAPDA. Extensive experiments demonstrate that AssistPDA outperforms existing offline VLM-based approaches, setting a new state-of-the-art for real-time VAPDA. Our dataset and code will be open-sourced to facilitate further research in the community.

📄 PDF Abstract BibTeX arXiv:2503.21904

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly DetectionAnomaly ForecastingQuestion AnsweringVideo Anomaly Detection

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding

2025-12-14 · Xinqi Jin, Hanxun Yu, Bohan Yu, Kebin Liu 외 arxiv

Online video understanding is essential for applications like public surveillance and AI glasses. However, applying Multimodal Large Language Models (MLLMs) to this domain is challenging due to the large number of video …

LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant

2025-03-05 · CVPR 2025 1 · Wei Li, Bing Hu, Rui Shao, Leyang Shen 외

First-person video assistants are highly anticipated to enhance our daily lives through online video dialogue. However, existing online video assistants often sacrifice assistant efficacy for real-time efficiency by proc…

Response Generation

Continual Learning for Anomaly Detection in Surveillance Videos

2020-04-15 · Keval Doshi, Yasin Yilmaz

Anomaly detection in surveillance videos has been recently gaining attention. A challenging aspect of high-dimensional applications such as video surveillance is continual learning. While current state-of-the-art deep le…

Anomaly DetectionAnomaly Detection In Surveillance VideosContinual LearningDecision Making+1

Online Anomaly Detection in Surveillance Videos with Asymptotic Bounds on False Alarm Rate

2020-10-10 · Keval Doshi, Yasin Yilmaz

Anomaly detection in surveillance videos is attracting an increasing amount of attention. Despite the competitive performance of recent methods, they lack theoretical performance analysis, particularly due to the complex…

Anomaly DetectionAnomaly Detection In Surveillance VideosDecision MakingVideo Anomaly Detection

Video Rain/Snow Removal by Transformed Online Multiscale Convolutional Sparse Coding

2019-09-13 · Minghan Li, Xiangyong Cao, Qian Zhao, Lei Zhang 외

Video rain/snow removal from surveillance videos is an important task in the computer vision community since rain/snow existed in videos can severely degenerate the performance of many surveillance system. Various method…

Snow Removal