paper-with-me

홈 › Papers

GeoDetect: Geometric Adversarial Detection for VLPs

2026-07-16 · Afsaneh Hasanebrahimi, Hanxun Huang, Christopher Leckie, James Bailey, Sarah Erfani arxiv

Vision-language pre-trained models (VLPs) are widely used in real-world applications. However, they remain vulnerable to adversarial attacks. Although adversarial detection methods have demonstrated success in single-modality settings (either vision or language), their effectiveness and reliability in multimodal models such as VLPs remain largely unexplored. In this work, we study the geometry of VLP embedding spaces and observe structured anisotropy that differs from unimodal vision models. Our theoretical analysis shows that under this anisotropic structure, adversarial attacks increase the expected geometric separation between clean and adversarial examples (AEs). Specifically, we demonstrate that AEs consistently exhibit greater expected distances to randomly sampled points than their clean counterparts, indicating that AEs tend to push representations out of manifold regions. Building on these insights, we propose GeoDetect, which leverages these off-manifold deviations via geometric scores to identify AEs. Through comprehensive evaluations, we show that our approach reliably detects AEs across diverse VLP architectures and threat settings, covering unimodal and multimodal attacks as well as adaptive attacks, thereby providing a robust and practical approach to improving the safety and reliability of these models.

📄 PDF Abstract BibTeX arXiv:2607.14737

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 111

Similar Papers 제목 키워드 기반

Watermarking Vision-Language Pre-trained Models for Multi-modal Embedding as a Service

2023-11-10 · Yuanmin Tang, Jing Yu, Keke Gai, Xiangyan Qu 외

Recent advances in vision-language pre-trained models (VLPs) have significantly increased visual understanding and cross-modal analysis capabilities. Companies have emerged to provide multi-modal Embedding as a Service (…

Model extraction

SARS-CoV-2 Virus-Like Particles with Plasmonic Au Cores and S1-Spike Protein Coronas

2022-12-27 · Weronika Andrzejewska, Barbara Peplińska, Jagoda Litowczenko, Patryk Obstarczyk 외

The COVID-19 pandemic has stimulated the scientific world to intensify virus-related studies, aimed at the development of quick and safe ways of detecting viruses in human body, studying the virus-antibody and virus-cell…

Compressing And Debiasing Vision-Language Pre-Trained Models for Visual Question Answering

2022-10-26 · Qingyi Si, Yuanxin Liu, Zheng Lin, Peng Fu 외

Despite the excellent performance of vision-language pre-trained models (VLPs) on conventional VQA task, they still suffer from two problems: First, VLPs tend to rely on language biases in datasets and fail to generalize…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Defense-Prefix for Preventing Typographic Attacks on CLIP

2023-04-10 · Hiroki Azuma, Yusuke Matsui

Vision-language pre-training models (VLPs) have exhibited revolutionary improvements in various vision-language tasks. In VLP, some adversarial attacks fool a model into false or absurd classifications. Previous studies …

object-detectionObject Detection

The Security Threat of Compressed Projectors in Large Vision-Language Models

2025-05-31 · Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen 외

The choice of a suitable visual language projector (VLP) is critical to the successful training of large visual language models (LVLMs). Mainstream VLPs can be broadly categorized into compressed and uncompressed project…

Computational Efficiency