paper-with-me

홈 › Papers

Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection

2026-02-27 · Zhaolin Cai, Fan Li, Huiyu Duan, Lijun He, Guangtao Zhai arxiv

Video anomaly detection (VAD) aims to identify abnormal events in videos. Traditional VAD methods generally suffer from the high costs of labeled data and full training, thus some recent works have explored leveraging frozen multi-modal large language models (MLLMs) in a tuning-free manner to perform VAD. However, their performance is limited as they directly inherit pre-training biases and cannot adapt internal representations to specific video contexts, leading to difficulties in handling subtle or ambiguous anomalies. To address these limitations, we propose a novel intervention framework, termed SteerVAD, which advances MLLM-based VAD by shifting from passively reading to actively steering and rectifying internal representations. Our approach first leverages the gradient-free representational separability analysis (RSA) to identify top attention heads as latent anomaly experts (LAEs) which are most discriminative for VAD. Then a hierarchical meta-controller (HMC) generates dynamic rectification signals by jointly conditioning on global context and these LAE outputs. The signals execute targeted, anisotropic scaling directly upon the LAE representation manifolds, amplifying anomaly-relevant dimensions while suppressing inherent biases. Extensive experiments on mainstream benchmarks demonstrate our method achieves state-of-the-art performance among tuning-free approaches requiring only 1% of training data, establishing it as a powerful new direction for video anomaly detection. The code will be released upon the publication.

📄 PDF Abstract BibTeX arXiv:2602.24021

Code (0)

등록된 구현이 없습니다.

Tasks

Video Anomaly Detection

Similar Papers 제목 키워드 기반

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

2026-07-22 · Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja 외 arxiv

Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refu…

Sparse Autoencoders as a Steering Basis for Phase Synchronization in Graph-Based CFD Surrogates

2026-03-28 · Yeping Hu, Ruben Glatt, Shusen Liu arxiv

Graph-based surrogate models provide fast alternatives to high-fidelity CFD solvers, but their opaque latent spaces and limited controllability restrict use in safety-critical settings. A key failure mode in oscillatory …

Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior

2026-06-07 · Tuc Nguyen, Thai Le arxiv

Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors. Most existing methods compute a fi…

Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models

2026-04-09 · Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee arxiv

We present a feedforward graph architecture in which heterogeneous frozen large language models serve as computational nodes, communicating through a shared continuous latent space via learned linear projections. Buildin…

Manifold-tiling Localized Receptive Fields are Optimal in Similarity-preserving Neural Networks

2018-12-01 · NeurIPS 2018 12 · Anirvan Sengupta, Cengiz Pehlevan, Mariano Tepper, Alexander Genkin 외

Many neurons in the brain, such as place cells in the rodent hippocampus, have localized receptive fields, i.e., they respond to a small neighborhood of stimulus space. What is the functional significance of such represe…

Hippocampus