paper-with-me

홈 › Papers

Towards Surveillance Video-and-Language Understanding: New Dataset Baselines and Challenges

2024-01-01 · CVPR 2024 1 · Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu, Chen Chen, Jian Jin, Zhenzhen Jiao

Surveillance videos are important for public security. However current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding although they have obtained considerable performance. To address this issue we propose a new research direction of surveillance video-and-language understanding(VALU) and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset UCA (UCF-Crime Annotation) contains 23542 sentences with an average length of 20 words and its annotated videos are as long as 110.7 hours. Furthermore we benchmark SOTA models for four multimodal tasks on this newly created dataset which serve as new baselines for surveillance VALU. Through experiments we find that mainstream models used in previously public datasets perform poorly on surveillance video demonstrating new challenges in surveillance VALU. We also conducted experiments on multimodal anomaly detection. These results demonstrate that our multimodal surveillance learning can improve the performance of anomaly detection. All the experiments highlight the necessity of constructing this dataset to advance surveillance AI.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly Detection

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges

2023-09-25 · Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu 외

Surveillance videos are an essential component of daily life with various critical applications, particularly in public security. However, current surveillance video tasks mainly focus on classifying and localizing anoma…

Anomaly DetectionDense Video CaptioningVideo CaptioningVideo Understanding

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

2025-05-19 · Bo Liu, Pengfei Qiao, Minhan Ma, Xuange Zhang 외

Understanding surveillance video content remains a critical yet underexplored challenge in vision-language research, particularly due to its real-world complexity, irregular event dynamics, and safety-critical implicatio…

Causal InferenceDecision MakingLanguage ModelingLanguage Modelling+2

Zero-Shot Action Recognition in Surveillance Videos

2024-10-28 · Joao Pereira, Vasco Lopes, David Semedo, Joao Neves

The growing demand for surveillance in public spaces presents significant challenges due to the shortage of human resources. Current AI-based video surveillance systems heavily rely on core computer vision models that re…

Action RecognitionVideo UnderstandingZero-Shot Action Recognition

ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance

2026-03-24 · Hyojin Park, Yi Li, Janghoon Cho, Sungha Choi 외 arxiv

Despite decades of work, surveillance still struggles in searching and reasoning about specific targets across long, multi-camera videos. Existing methods - tracking, retrieval, and video LLMs require heavy manual filter…

Multimodal ReasoningQuestion Answering

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

2026-01-08 · Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera arxiv

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate …

Visual ReasoningVideo Retrieval