paper-with-me

홈 › Papers

PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding

2026-01-06 · Iñaki Erregue, Kamal Nasrollahi, Sergio Escalera arxiv

Video Anomaly Understanding (VAU) extends traditional Video Anomaly Detection (VAD) by not only localizing anomalies but also describing and reasoning about their context. Existing VAU approaches often rely on fine-tuned multimodal large language models (MLLMs) or external modules such as video captioners, which introduce costly annotations, complex training pipelines, and high inference overhead. In this work, we introduce PrismVAU, a lightweight yet effective system for real-time VAU that leverages a single off-the-shelf MLLM for anomaly scoring, explanation, and prompt optimization. PrismVAU operates in two complementary stages: (1) a coarse anomaly scoring module that computes frame-level anomaly scores via similarity to textual anchors, and (2) an MLLM-based refinement module that contextualizes anomalies through system and user prompts. Both textual anchors and prompts are optimized with a weakly supervised Automatic Prompt Engineering (APE) framework. Extensive experiments on standard VAD benchmarks demonstrate that PrismVAU delivers competitive detection performance and interpretable anomaly explanations -- without relying on instruction tuning, frame-level annotations, and external modules or dense processing -- making it an efficient and practical solution for real-world applications.

📄 PDF Abstract BibTeX arXiv:2601.02927

Code (0)

등록된 구현이 없습니다.

Tasks

Video Anomaly DetectionPrompt Engineering

Similar Papers 제목 키워드 기반

Open Multimodal Retrieval-Augmented Factual Image Generation

2025-10-26 · Yang Tian, Fan Liu, Jingyuan Zhang, Wei Bi 외 arxiv

Large Multimodal Models (LMMs) have achieved remarkable progress in generating photorealistic and prompt-aligned images, but they often produce outputs that contradict verifiable knowledge, especially when prompts involv…

Image Generation

Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis

2025-12-27 · Zhiqiang Gao, Shihao Gao, Zixing Zhang, Yihao Guo 외 arxiv

Understanding sentiment in multimodal conversations is a complex yet crucial challenge toward building emotionally intelligent AI systems. The Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) Challenge …

Multimodal Sentiment Analysis

PromptLoop: Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment

2025-10-01 · Suhyeon Lee, Jong Chul Ye arxiv

Despite recent progress, reinforcement learning (RL)-based fine-tuning of diffusion models often struggles with generalization, composability, and robustness against reward hacking. Recent studies have explored prompt re…

Reinforcement Learning

Prompting ChatGPT in MNER: Enhanced Multimodal Named Entity Recognition with Auxiliary Refined Knowledge

2023-05-20 · Jinyuan Li, Han Li, Zhuo Pan, Di Sun 외

Multimodal Named Entity Recognition (MNER) on social media aims to enhance textual entity prediction by incorporating image-based clues. Existing studies mainly focus on maximizing the utilization of pertinent image info…

Multi-modal Named Entity Recognitionnamed-entity-recognitionNamed Entity Recognition

Adversarial Prompt Injection Attack on Multimodal Large Language Models

2026-03-31 · Meiwen Ding, Song Xia, Chenqi Kong, Xudong Jiang arxiv

Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks. Existing prompt injection m…