paper-with-me

홈 › Papers

DeepSVU: Towards In-depth Security-oriented Video Understanding via Unified Physical-world Regularized MoE

2026-02-20 · Yujie Jin, Wenxin Zhang, Jingjing Wang, Guodong Zhou arxiv

In the literature, prior research on Security-oriented Video Understanding (SVU) has predominantly focused on detecting and localize the threats (e.g., shootings, robberies) in videos, while largely lacking the effective capability to generate and evaluate the threat causes. Motivated by these gaps, this paper introduces a new chat paradigm SVU task, i.e., In-depth Security-oriented Video Understanding (DeepSVU), which aims to not only identify and locate the threats but also attribute and evaluate the causes threatening segments. Furthermore, this paper reveals two key challenges in the proposed task: 1) how to effectively model the coarse-to-fine physical-world information (e.g., human behavior, object interactions and background context) to boost the DeepSVU task; and 2) how to adaptively trade off these factors. To tackle these challenges, this paper proposes a new Unified Physical-world Regularized MoE (UPRM) approach. Specifically, UPRM incorporates two key components: the Unified Physical-world Enhanced MoE (UPE) Block and the Physical-world Trade-off Regularizer (PTR), to address the above two challenges, respectively. Extensive experiments conduct on our DeepSVU instructions datasets (i.e., UCF-C instructions and CUVA instructions) demonstrate that UPRM outperforms several advanced Video-LLMs as well as non-VLM approaches. Such information.These justify the importance of the coarse-to-fine physical-world information in the DeepSVU task and demonstrate the effectiveness of our UPRM in capturing such information.

📄 PDF Abstract BibTeX arXiv:2602.18019

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spatial-Conditioned Reasoning in Long-Egocentric Videos

2026-01-26 · James Tribble, Hao Wang, Si-En Hong, Chaoyi Zhou 외 arxiv

Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and…

Spatial ReasoningVisual Navigation

Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding

2025-05-19 · Thong Nguyen, Zhiyuan Hu, Xu Lin, Cong-Duy Nguyen 외

Recent years have witnessed outstanding advances of large vision-language models (LVLMs). In order to tackle video understanding, most of them depend upon their implicit temporal understanding capacity. As such, they hav…

Language ModelingLanguage ModellingLarge Language ModelVideo Understanding

SNEAK: Synonymous Sentences-Aware Adversarial Attack on Natural Language Video Localization

2021-12-08 · Wenbo Gou, Wen Shi, Jian Lou, Lijie Huang 외

Natural language video localization (NLVL) is an important task in the vision-language understanding area, which calls for an in-depth understanding of not only computer vision and natural language side alone, but more i…

Adversarial AttackAdversarial Robustness

EgoTaskQA: Understanding Human Tasks in Egocentric Videos

2022-10-08 · Baoxiong Jia, Ting Lei, Song-Chun Zhu, Siyuan Huang

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, t…

Action LocalizationcounterfactualDescriptiveDiagnostic+3

Long Activity Video Understanding using Functional Object-Oriented Network

2018-07-03 · Ahmad Babaeian Jelodar, David Paulius, Yu Sun

Video understanding is one of the most challenging topics in computer vision. In this paper, a four-stage video understanding pipeline is presented to simultaneously recognize all atomic actions and the single on-going a…

ObjectVideo Understanding