Consistency-based Self-supervised Learning for Temporal Anomaly Localization
This work tackles Weakly Supervised Anomaly detection, in which a predictor is allowed to learn not only from normal examples but also from a few labeled anomalies made available during training. In particular, we deal with the localization of anomalous activities within the video stream: this is a very challenging scenario, as training examples come only with video-level annotations (and not frame-level). Several recent works have proposed various regularization terms to address it i.e. by enforcing sparsity and smoothness constraints over the weakly-learned frame-level anomaly scores. In this work, we get inspired by recent advances within the field of self-supervised learning and ask the model to yield the same scores for different augmentations of the same video sequence. We show that enforcing such an alignment improves the performance of the model on XD-Violence.
Code (1)
Tasks
Anomaly Detection In Surveillance VideosAnomaly LocalizationSelf-Supervised LearningSupervised Anomaly DetectionWeakly-supervised Anomaly DetectionWeakly-supervised Temporal Action LocalizationSimilar Papers 제목 키워드 기반
Learning Where and When: Patch-Based Spatiotemporal Localization in Weakly Supervised Video Anomaly Detection
Weakly supervised video anomaly detection (WSVAD) has predominantly focused on temporal localization, identifying when anomalies occur while largely neglecting their spatial extent within frames. Yet, spatial localizatio…
Multiple Instance LearningVideo Anomaly DetectionSelf-Supervised Masking for Unsupervised Anomaly Detection and Localization
Recently, anomaly detection and localization in multimedia data have received significant attention among the machine learning community. In real-world applications such as medical diagnosis and industrial defect detecti…
Anomaly DetectionAnomaly LocalizationDefect DetectionMedical Diagnosis+2Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
Point-supervised Temporal Action Localization (PTAL) adopts a lightly frame-annotated paradigm (\textit{i.e.}, labeling only a single frame per action instance) to train a model to effectively locate action instances wit…
Weakly-supervised Temporal Action LocalizationMulti-Task LearningWhat's the Catch? Evaluating Temporal Consistency in Vision-Language Models
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as …
Anomaly DetectionDenoising Diffusion Models for Anomaly Localization in Medical Images
This chapter explores anomaly localization in medical images using denoising diffusion models. After providing a brief methodological background of these models, including their application to image reconstruction and th…
Anomaly LocalizationDenoisingImage Reconstruction