paper-with-me

홈 › Papers

Video Monitoring Queries

2020-02-24 · Nick Koudas, Raymond Li, Ioannis Xarchakos

Recent advances in video processing utilizing deep learning primitives achieved breakthroughs in fundamental problems in video analysis such as frame classification and object detection enabling an array of new applications. In this paper we study the problem of interactive declarative query processing on video streams. In particular we introduce a set of approximate filters to speed up queries that involve objects of specific type (e.g., cars, trucks, etc.) on video frames with associated spatial relationships among them (e.g., car left of truck). The resulting filters are able to assess quickly if the query predicates are true to proceed with further analysis of the frame or otherwise not consider the frame further avoiding costly object detection operations. We propose two classes of filters $IC$ and $OD$, that adapt principles from deep image classification and object detection. The filters utilize extensible deep neural architectures and are easy to deploy and utilize. In addition, we propose statistical query processing techniques to process aggregate queries involving objects with spatial constraints on video streams and demonstrate experimentally the resulting increased accuracy on the resulting aggregate estimation. Combined these techniques constitute a robust set of video monitoring query processing techniques. We demonstrate that the application of the techniques proposed in conjunction with declarative queries on video streams can dramatically increase the frame processing rate and speed up query processing by at least two orders of magnitude. We present the results of a thorough experimental study utilizing benchmark video data sets at scale demonstrating the performance benefits and the practical relevance of our proposals.

📄 PDF Abstract BibTeX arXiv:2002.10537

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage Classificationobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Eyes on the Road: State-of-the-Art Video Question Answering Models Assessment for Traffic Monitoring Tasks

2024-12-02 · Joseph Raj Vishal, Divesh Basina, Aarya Choudhary, Bharatesh Chakravarthi

Recent advances in video question answering (VideoQA) offer promising applications, especially in traffic monitoring, where efficient video interpretation is critical. Within ITS, answering complex, real-time queries lik…

Multi-Object TrackingObject TrackingQuestion AnsweringVideo Question Answering

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

2026-05-09 · Tri Cao, Khoi Le, Thong Nguyen, Cong-Duy Nguyen 외 arxiv

While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the abilit…

QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon Tracking

2026-05-10 · Mayank Anand, Mohammad Saqlain, Kyan Mahajan, Priya Shukla 외 arxiv

Tracking points in videos is typically formulated as frame-to-frame correspondence, where each point is matched locally to the next frame. While this works over short horizons, errors accumulate under articulation, occlu…

I-ViSE: Interactive Video Surveillance as an Edge Service using Unsupervised Feature Queries

2020-03-09 · Seyed Yahya Nikouei, Yu Chen, Alexander Aved, Erik Blasch

Situation AWareness (SAW) is essential for many mission critical applications. However, SAW is very challenging when trying to immediately identify objects of interest or zoom in on suspicious activities from thousands o…

DescriptiveFace RecognitionScene Recognition

UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation

2026-07-09 · Mingyu Dou, Shi Qiu, Ming Hu, Yifan Chen 외 arxiv

Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video perception is largely limited to box-level detection and tracking over prede…

Video Instance Segmentation