paper-with-me

홈 › Papers

EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy

2026-05-15 · Xuanyu Ge, Zhongqi Wang, Jie Zhang, Shiguang Shan, Xilin Chen arxiv

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to backdoor attacks. Existing defense methods predominantly focus on sample-level defense, which relies on the knowledge of training data or triggers. However, identifying whether a given model is backdoored remains a critical but unexplored task. To fill this gap, we propose EntropyScan, a lightweight and trigger-agnostic method for model-level backdoor detection in LVLMs. We first observe that backdoor injection disrupts the cross-modal alignment, resulting in pronounced structural anomalies in visual attention allocation on benign samples. Based on this insight, EntropyScan detects the backdoor models by quantifying such attention deviations. Specifically, it extracts visual attention distributions from the initial layers of the Large Language Model (LLM) and applies Tsallis entropy to capture these structural distortions. By employing a reference-anchored Z-score normalization on a small set of benign samples, it effectively identifies the backdoored model. Extensive experiments across two LVLMs architectures and three advanced attack scenarios show that EntropyScan achieves an F1 score of 98.5% in average and an AUC of 96.6%. Our code will be publicly available soon.

📄 PDF Abstract BibTeX arXiv:2605.15711

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Models

2025-02-25 · CVPR 2025 1 · Zhaoyi Liu, huan zhang

Self-supervised learning (SSL) vision encoders learn high-quality image representations and thus have become a vital part of developing vision modality of large vision language models (LVLMs). Due to the high cost of tra…

Backdoor AttackHallucinationSelf-Supervised Learning

Test-Time Attention Purification for Backdoored Large Vision Language Models

2026-03-13 · Zhifang Zhang, Bojun Yang, Shuo He, Weitong Chen 외 arxiv

Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded samples into the training data to implant b…

Robust Anti-Backdoor Instruction Tuning in LVLMs

2025-06-04 · Yuan Xun, Siyuan Liang, Xiaojun Jia, Xinwei Liu 외

Large visual language models (LVLMs) have demonstrated excellent instruction-following capabilities, yet remain vulnerable to stealthy backdoor attacks when finetuned using contaminated data. Existing backdoor defense te…

backdoor defenseInstruction Following

MTAttack: Multi-Target Backdoor Attacks against Large Vision-Language Models

2025-11-13 · Zihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng 외 arxiv

Recent advances in Large Visual Language Models (LVLMs) have demonstrated impressive performance across various vision-language tasks by leveraging large-scale image-text pretraining and instruction tuning. However, the …

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models

2025-09-29 · Zhifang Zhang, Qiqi Tao, Jiaqi Lv, Na Zhao 외 arxiv

Large vision-language models (LVLMs) have achieved impressive performance across a wide range of vision-language tasks, while they remain vulnerable to backdoor attacks. Existing backdoor attacks on LVLMs aim to force th…