paper-with-me

홈 › Papers

OneFocus: Enabling Real-World X-ray Security Screening with a Unified Vision-Language Model

2026-06-14 · Jiali Wen, Hongxia Gao, Litao Li, Yixin Chen, Kaijie Zhang, Qianyun Liu, Xiaoqin Wen arxiv

X-ray contraband detection is critical for security in large-scale logistics and transportation, yet conventional detectors struggle to adapt to emerging contraband types and lack fundamental visual understanding. Vision-language models (VLMs) offer strong generalization but are hindered by the scarcity of high-quality X-ray image-caption data. To bridge this critical gap, we present MMXray, a meticulously curated benchmark of 52,124 image-caption pairs spanning 28 fine-grained classes of X-ray contraband. To enrich MMXray with realistic occlusion patterns, we further introduce CleanDET, a dedicated synthesis dataset containing clean foreground contraband images from 28 categories and background images with diverse density levels, together with AnyContraSyn, a controllable synthesis method designed to operate on CleanDET. We also develop OnePipe, an extensible pipeline for systematic data curation. Built on MMXray, we propose OneFocus, a unified VLM that supports four core tasks: visual question answering, contraband localization, classification, and image understanding. OneFocus achieves state-of-the-art performance in X-ray contraband understanding and demonstrates robust cross-domain generalization, establishing a strong vision-language baseline for security screening.

📄 PDF Abstract BibTeX arXiv:2606.15663

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringDomain Generalization

Similar Papers 제목 키워드 기반

STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection

2025-04-03 · CVPR 2025 1 · Divya Velayudhan, Abdelfatah Ahmed, Mohamad Alansari, Neha Gour 외

Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated…

Instruction FollowingLanguage ModelingLanguage ModellingQuestion Answering+3

CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

2026-06-03 · Tianneng Shi, Robin Rheem, Dongwei Jiang, Mona Wang 외 arxiv

AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabilities. However, existing cybersecurity evaluations of AI systems are limited in …

Multi-Class 3D Object Detection Within Volumetric 3D Computed Tomography Baggage Security Screening Imagery

2020-08-03 · Qian Wang, Neelanjan Bhowmik, Toby P. Breckon

Automatic detection of prohibited objects within passenger baggage is important for aviation security. X-ray Computed Tomography (CT) based 3D imaging is widely used in airports for aviation security screening whilst pri…

3D Object DetectionComputed Tomography (CT)Data AugmentationObject+2

Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios

2026-07-25 · Lixun Ma, Ruolong Ma, Bei Wang, Feng Wei 외 arxiv

Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains underexplored. Existing benchmarks often rely on explicitly specified security requ…

Code Generation

On the Evaluation of Prohibited Item Classification and Detection in Volumetric 3D Computed Tomography Baggage Security Screening Imagery

2020-03-27 · Qian Wang, Neelanjan Bhowmik, Toby P. Breckon

X-ray Computed Tomography (CT) based 3D imaging is widely used in airports for aviation security screening whilst prior work on prohibited item detection focuses primarily on 2D X-ray imagery. In this paper, we aim to ev…

3D Object DetectionComputed Tomography (CT)General ClassificationObject+2