The Security Threat of Compressed Projectors in Large Vision-Language Models
The choice of a suitable visual language projector (VLP) is critical to the successful training of large visual language models (LVLMs). Mainstream VLPs can be broadly categorized into compressed and uncompressed projectors, and each offering distinct advantages in performance and computational efficiency. However, their security implications have not been thoroughly examined. Our comprehensive evaluation reveals significant differences in their security profiles: compressed projectors exhibit substantial vulnerabilities, allowing adversaries to successfully compromise LVLMs even with minimal knowledge of structural information. In stark contrast, uncompressed projectors demonstrate robust security properties and do not introduce additional vulnerabilities. These findings provide critical guidance for researchers in selecting optimal VLPs that enhance the security and reliability of visual language models. The code will be released.
Code (0)
등록된 구현이 없습니다.
Tasks
Computational EfficiencySimilar Papers 제목 키워드 기반
Bitstream Collisions in Neural Image Compression via Adversarial Perturbations
Neural image compression (NIC) has emerged as a promising alternative to classical compression techniques, offering improved compression ratios. Despite its progress towards standardization and practical deployment, ther…
Adversarial AttackImage CompressionDeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
The visual projector, which bridges the vision and language modalities and facilitates cross-modal alignment, serves as a crucial component in MLLMs. However, measuring the effectiveness of projectors in vision-language …
cross-modal alignmentVisual LocalizationVisual Question Answering (VQA)STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated…
Instruction FollowingLanguage ModelingLanguage ModellingQuestion Answering+3A new GAN-based anomaly detection (GBAD) approach for multi-threat object classification on large-scale x-ray security images
Threat object recognition in x-ray security images is one of the important practical applications of computer vision. However, research in this field has been limited by the lack of available dataset that would mirror th…
Anomaly DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONObject RecognitionEvaluating GAN-Based Image Augmentation for Threat Detection in Large-Scale Xray Security Images
The inherent imbalance in the data distribution of X-ray security images is one of the most challenging aspects of computer vision algorithms applied in this domain. Most of the prior studies in this field have ignored t…
Image AugmentationImage GenerationObject Detection