paper-with-me

홈 › Papers

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought

2025-05-26 · Chao Huang, Benfeng Wang, Jie Wen, Chengliang Liu, Wei Wang, Li Shen, Xiaochun Cao

Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLM to reason anomaly step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks. Codes and datasets will be released at https://github.com/wbfwonderful/Vad-R1.

📄 PDF Abstract BibTeX arXiv:2505.19877

Code (1)

wbfwonderful/vad-r1 공식 구현

Tasks

Anomaly DetectionVideo Anomaly Detection

Similar Papers 제목 키워드 기반

Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method

2026-01-15 · Chao Huang, Benfeng Wang, Wei Wang, Jie Wen 외 arxiv

Recent progress in reasoning capabilities of Multimodal Large Language Models(MLLMs) has highlighted their potential for performing complex video understanding tasks. However, in the domain of Video Anomaly Detection and…

Video Anomaly DetectionDecision Making

VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning

2025-05-29 · Liyun Zhu, Qixiang Chen, Xi Shen, Xiaodong Cun

Video Anomaly Understanding (VAU) is essential for applications such as smart cities, security surveillance, and disaster alert systems, yet remains challenging due to its demand for fine-grained spatio-temporal percepti…

Anomaly DetectionDescriptiveMultiple-choiceQuestion Answering

Chain-of-Anomaly Thoughts with Large Vision-Language Models

2025-12-23 · Pedro Domingos, João Pereira, Vasco Lopes, João Neves 외 arxiv

Automated video surveillance with Large Vision-Language Models is limited by their inherent bias towards normality, often failing to detect crimes. While Chain-of-Thought reasoning strategies show significant potential f…

Anomaly ClassificationAnomaly Detection

A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis

2025-11-02 · Dongheng Lin, Mengxue Qu, Kunyang Han, Jianbo Jiao 외 arxiv

Most video-anomaly research stops at frame-wise detection, offering little insight into why an event is abnormal, typically outputting only frame-wise anomaly scores without spatial or semantic context. Recent video anom…

Video Anomaly Detection

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

2026-05-11 · Xiran Zhao, Jing Jin, Yan Bai, Zhongan Wang 외 arxiv

Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which do not fully reflect continuous inspection processes in real industri…

Anomaly Detection