paper-with-me

홈 › Papers

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

2025-05-19 · Bo Liu, Pengfei Qiao, Minhan Ma, Xuange Zhang, Yinan Tang, Peng Xu, Kun Liu, Tongtong Yuan

Understanding surveillance video content remains a critical yet underexplored challenge in vision-language research, particularly due to its real-world complexity, irregular event dynamics, and safety-critical implications. In this work, we introduce SurveillanceVQA-589K, the largest open-ended video question answering benchmark tailored to the surveillance domain. The dataset comprises 589,380 QA pairs spanning 12 cognitively diverse question types, including temporal reasoning, causal inference, spatial understanding, and anomaly interpretation, across both normal and abnormal video scenarios. To construct the benchmark at scale, we design a hybrid annotation pipeline that combines temporally aligned human-written captions with Large Vision-Language Model-assisted QA generation using prompt-based techniques. We also propose a multi-dimensional evaluation protocol to assess contextual, temporal, and causal comprehension. We evaluate eight LVLMs under this framework, revealing significant performance gaps, especially in causal and anomaly-related tasks, underscoring the limitations of current models in real-world surveillance contexts. Our benchmark provides a practical and comprehensive resource for advancing video-language understanding in safety-critical applications such as intelligent monitoring, incident analysis, and autonomous decision-making.

📄 PDF Abstract BibTeX arXiv:2505.12589

Code (0)

등록된 구현이 없습니다.

Tasks

Causal InferenceDecision MakingLanguage ModelingLanguage ModellingQuestion AnsweringVideo Question Answering

Similar Papers 제목 키워드 기반

Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges

2023-09-25 · Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu 외

Surveillance videos are an essential component of daily life with various critical applications, particularly in public security. However, current surveillance video tasks mainly focus on classifying and localizing anoma…

Anomaly DetectionDense Video CaptioningVideo CaptioningVideo Understanding

Towards Surveillance Video-and-Language Understanding: New Dataset Baselines and Challenges

2024-01-01 · CVPR 2024 1 · Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu 외

Surveillance videos are important for public security. However current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying t…

Anomaly Detection

SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models

2026-01-08 · Oriol Rabasseda, Zenjie Li, Kamal Nasrollahi, Sergio Escalera arxiv

Automatic identification of events and recurrent behavior analysis are critical for video surveillance. However, most existing content-based video retrieval benchmarks focus on scene-level similarity and do not evaluate …

Visual ReasoningVideo Retrieval

UAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark

2024-10-02 · Hasnat Md Abdullah, Tian Liu, Kangda Wei, Shu Kong 외

Localizing unusual activities, such as human errors or surveillance incidents, in videos holds practical significance. However, current video understanding models struggle with localizing these unusual events likely beca…

Unusual Activity LocalizationVideo Understanding

IJB–S: IARPA Janus Surveillance Video Benchmark

2019-04-25 · Nathan D. Kalka Noblis, Brianna Maze; James A. Duncan, Kevin O’Connor, Stephen Elliott 외

We present IJB-S dataset, an open-source IARPA Janus Surveillance Video Benchmark and associated protocols. The dataset consists of images and surveillance video collected from 202 subjects at a Department of Defense (Do…