paper-with-me

홈 › Papers

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders

2025-05-30 · Yuqi Zhang, Yuchun Miao, Zuchao Li, Liang Ding

We introduce AMIA, a lightweight, inference-only defense for Large Vision-Language Models (LVLMs) that (1) Automatically Masks a small set of text-irrelevant image patches to disrupt adversarial perturbations, and (2) conducts joint Intention Analysis to uncover and mitigate hidden harmful intents before response generation. Without any retraining, AMIA improves defense success rates across diverse LVLMs and jailbreak benchmarks from an average of 52.4% to 81.7%, preserves general utility with only a 2% average accuracy drop, and incurs only modest inference overhead. Ablation confirms both masking and intention analysis are essential for a robust safety-utility trade-off.

📄 PDF Abstract BibTeX arXiv:2505.24519

Code (0)

등록된 구현이 없습니다.

Tasks

Response Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ModalImmune: Immunity Driven Unlearning via Self Destructive Training

2026-02-18 · Rong Fu, WeiZhi Tang, Ziming Wang, Jia Yee Tan 외 arxiv

Multimodal systems are vulnerable to partial or complete loss of input channels at deployment, which undermines reliability in real-world settings. This paper presents ModalImmune, a training framework that enforces moda…

Resolving issues of scaling for gramian based input-output pairing methods

2019-10-22

A key problem in process control is to decide which inputs should control which outputs. There are multiple ways to solve this problem, among them using gramian based measures, which include the Hankel interaction index …

Automatic Identification of Cuneiform Fragments Using String Alignment Algorithms

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The literature from ancient Mesopotamia is still riddled with textual lacunas. Scores of fragments which could potentially fill those lacunas lie unidentified in museums's cabinets, but their identification has tradition…

PANORAMIA: Privacy Auditing of Machine Learning Models without Retraining

2024-02-12 · Mishaal Kazmi, Hadrien Lautraite, Alireza Akbari, Qiaoyue Tang 외

We present PANORAMIA, a privacy leakage measurement framework for machine learning models that relies on membership inference attacks using generated data as non-members. By relying on generated non-member data, PANORAMI…

Robust Automatic Differentiation of Square-Root Kalman Filters via Gramian Differentials

2026-03-13 · Adrien Corenflos arxiv

Square-root Kalman filters propagate state covariances in Cholesky-factor form for numerical stability, and are a natural target for gradient-based parameter learning in state-space models. Their core operation, triangul…