Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
Recent weakly supervised video anomaly detection methods have achieved significant advances by employing unified frameworks for joint optimization. However, this paradigm is limited by a fundamental sensitivity-stability trade-off, as the conflicting objectives for detecting transient and sustained anomalies lead to either fragmented predictions or over-smoothed responses. To address this limitation, we propose DeSC, a novel Decoupled Sensitivity-Consistency framework that trains two specialized streams using distinct optimization strategies. The temporal sensitivity stream adopts an aggressive optimization strategy to capture high-frequency abrupt changes, whereas the semantic consistency stream applies robust constraints to maintain long-term coherence and reduce noise. Their complementary strengths are fused through a collaborative inference mechanism that reduces individual biases and produces balanced predictions. Extensive experiments demonstrate that DeSC establishes new state-of-the-art performance by achieving 89.37% AUC on UCF-Crime (+1.29%) and 87.18% AP on XD-Violence (+2.22%). Code is available at https://github.com/imzht/DeSC.
Code (0)
등록된 구현이 없습니다.
Tasks
Video Anomaly DetectionResults from the Paper
| Rank | Task | Dataset | Model | Metrics |
|---|---|---|---|---|
| #5 | Anomaly Detection | UCF-Crime | DeSC | AUC: 89.37 |
| #5 | Video Anomaly Detection | UCF-Crime | DeSC | AUC: 89.37 |
Similar Papers 제목 키워드 기반
Weakly-supervised Micro- and Macro-expression Spotting Based on Multi-level Consistency
Most micro- and macro-expression spotting methods in untrimmed videos suffer from the burden of video-wise collection and frame-wise annotation. Weakly-supervised expression spotting (WES) based on video-level labels can…
Multiple Instance LearningOptical Flow EstimationDual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning
Unsupervised video object-centric learning aims to decompose dynamic scenes into persistent, object-level representations without supervision. However, existing slot-based methods struggle to maintain stable object ident…
Object RecognitionFast Weakly Supervised Action Segmentation Using Mutual Consistency
Action segmentation is the task of predicting the actions for each frame of a video. As obtaining the full annotation of videos for action segmentation is expensive, weakly supervised approaches that can learn only from …
Action SegmentationSegmentationWeakly Supervised Action Segmentation (Transcript)Weakly Supervised Instance Segmentation for Videos with Temporal Mask Consistency
Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) parti…
Instance SegmentationRelation NetworkSegmentationSemantic Segmentation+1Decoupled Spatial Neural Attention for Weakly Supervised Semantic Segmentation
Weakly supervised semantic segmentation receives much research attention since it alleviates the need to obtain a large amount of dense pixel-wise ground-truth annotations for the training images. Compared with other for…
Image CaptioningSegmentationSemantic SegmentationWeakly supervised Semantic Segmentation+1