paper-with-me

Papers

The Security Threat of Compressed Projectors in Large Vision-Language Models

2025-05-31 · Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Zhanhui Kang, Di Wang, Yu Wang

The choice of a suitable visual language projector (VLP) is critical to the successful training of large visual language models (LVLMs). Mainstream VLPs can be broadly categorized into compressed and uncompressed projectors, and each offering distinct advantages in performance and computational efficiency. However, their security implications have not been thoroughly examined. Our comprehensive evaluation reveals significant differences in their security profiles: compressed projectors exhibit substantial vulnerabilities, allowing adversaries to successfully compromise LVLMs even with minimal knowledge of structural information. In stark contrast, uncompressed projectors demonstrate robust security properties and do not introduce additional vulnerabilities. These findings provide critical guidance for researchers in selecting optimal VLPs that enhance the security and reliability of visual language models. The code will be released.

📄 PDF Abstract BibTeX arXiv:2506.00534

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Bitstream Collisions in Neural Image Compression via Adversarial Perturbations

2025-03-25 · Jordan Madden, Lhamo Dorje, Xiaohua LI

Neural image compression (NIC) has emerged as a promising alternative to classical compression techniques, offering improved compression ratios. Despite its progress towards standardization and practical deployment, ther…

Adversarial AttackImage Compression

DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

2024-05-31 · Linli Yao, Lei LI, Shuhuai Ren, Lean Wang 외

The visual projector, which bridges the vision and language modalities and facilitates cross-modal alignment, serves as a crucial component in MLLMs. However, measuring the effectiveness of projectors in vision-language …

cross-modal alignmentVisual LocalizationVisual Question Answering (VQA)

STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection

2025-04-03 · CVPR 2025 1 · Divya Velayudhan, Abdelfatah Ahmed, Mohamad Alansari, Neha Gour 외

Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated…

Instruction FollowingLanguage ModelingLanguage ModellingQuestion Answering+3

A new GAN-based anomaly detection (GBAD) approach for multi-threat object classification on large-scale x-ray security images

2019-10-23 · IEICE Transactions on Information and Systems 2019 10 · Joanna Kazzandra Dumagpi, Woo-Young Jung, Yong-Jin Jeong

Threat object recognition in x-ray security images is one of the important practical applications of computer vision. However, research in this field has been limited by the lack of available dataset that would mirror th…

Anomaly DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONObject Recognition

Evaluating GAN-Based Image Augmentation for Threat Detection in Large-Scale Xray Security Images

2020-12-23 · Applied Sciences 2020 12 · Joanna Kazzandra Dumagpi, Yong-Jin Jeong

The inherent imbalance in the data distribution of X-ray security images is one of the most challenging aspects of computer vision algorithms applied in this domain. Most of the prior studies in this field have ignored t…

Image AugmentationImage GenerationObject Detection