paper-with-me

홈 › Papers

Transformer-Driven Multimodal Fusion for Explainable Suspiciousness Estimation in Visual Surveillance

2025-12-10 · Kuldeep Singh Yadav, Lalan Kumar arxiv

Suspiciousness estimation is critical for proactive threat detection and ensuring public safety in complex environments. This work introduces a large-scale annotated dataset, USE50k, along with a computationally efficient vision-based framework for real-time suspiciousness analysis. The USE50k dataset contains 65,500 images captured from diverse and uncontrolled environments, such as airports, railway stations, restaurants, parks, and other public areas, covering a broad spectrum of cues including weapons, fire, crowd density, abnormal facial expressions, and unusual body postures. Building on this dataset, we present DeepUSEvision, a lightweight and modular system integrating three key components, i.e., a Suspicious Object Detector based on an enhanced YOLOv12 architecture, dual Deep Convolutional Neural Networks (DCNN-I and DCNN-II) for facial expression and body-language recognition using image and landmark features, and a transformer-based Discriminator Network that adaptively fuses multimodal outputs to yield an interpretable suspiciousness score. Extensive experiments confirm the superior accuracy, robustness, and interpretability of the proposed framework compared to state-of-the-art approaches. Collectively, the USE50k dataset and the DeepUSEvision framework establish a strong and scalable foundation for intelligent surveillance and real-time risk assessment in safety-critical applications.

📄 PDF Abstract BibTeX arXiv:2512.09311

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Advancing Multimodal Data Fusion in Pain Recognition: A Strategy Leveraging Statistical Correlation and Human-Centered Perspectives

2024-03-30

This research presents a novel multimodal data fusion methodology for pain behavior recognition, integrating statistical correlation analysis with human-centered insights. Our approach introduces two key innovations: 1) …

Exploring the In-Context Learning Capabilities of LLMs for Money Laundering Detection in Financial Graphs

2025-07-20 · Erfan Pirmorad arxiv

The complexity and interconnectivity of entities involved in money laundering demand investigative reasoning over graph-structured data. This paper explores the use of large language models (LLMs) as reasoning engines ov…

Graded Suspiciousness of Adversarial Texts to Human

2024-10-06 · Shakila Mahjabin Tonni, Pedro Faustini, Mark Dras

Adversarial examples pose a significant challenge to deep neural networks (DNNs) across both image and text domains, with the intent to degrade model performance through meticulously altered inputs. Adversarial texts, ho…

Adversarial AttackAdversarial TextSemantic SimilaritySemantic Textual Similarity+1

A Fully Transformer Based Multimodal Framework for Explainable Cancer Image Segmentation Using Radiology Reports

2025-08-19 · Enobong Adahada, Isabel Sassoon, Kate Hone, Yongmin Li arxiv

We introduce Med-CTX, a fully transformer based multimodal framework for explainable breast cancer ultrasound segmentation. We integrate clinical radiology reports to boost both performance and interpretability. Med-CTX …

Image Segmentation

IsoNet: Causal Analysis of Multimodal Transformers for Neuromuscular Gesture Classification

2025-06-20 · Eion Tyacke, Kunal Gupta, Jay Patel, Rui Li

Hand gestures are a primary output of the human motor system, yet the decoding of their neuromuscular signatures remains a bottleneck for basic neuroscience and assistive technologies such as prosthetics. Traditional hum…