paper-with-me

홈 › Papers

When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents

2026-06-22 · Yanhang Li, Zhichao Fan, Zexin Zhuang arxiv

Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for flagging indirect prompt injection (IPI) in multimodal computer-use agents before the agent emits a corrupted action. We argue, on a single-backbone cautionary case study (Qwen2.5-VL-7B on Mind2Web, teacher-forced replay), that a high probing AUC on a clean-vs-attack split is not, on its own, evidence of malicious-content detection. Two post-hoc diagnostics -- a paired-construction scalar baseline on text-side injections, and same-step nuisance-matched visual controls on the overlay surface -- do not license an unqualified malicious-content interpretation of the headline while leaving room for partly-semantic readings. We package the diagnostics as a candidate control set with reporting heuristics for what a high clean-vs-attack AUC does and does not license. Labels are injection-surface-present, not attack success; generalisation beyond this backbone and benchmark is a conjecture.

📄 PDF Abstract BibTeX arXiv:2606.22864

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Good-Enough Example Extrapolation

2021-09-12 · EMNLP 2021 11 · Jason Wei

This paper asks whether extrapolating the hidden space distribution of text examples from one class onto another is a valid inductive bias for data augmentation. To operationalize this question, I propose a simple data a…

Data AugmentationInductive Biastext-classificationText Classification+1

Visual Dialog

2016-11-26 · CVPR 2017 7 · Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh 외

We introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a q…

AI AgentChatbotRetrievalVisual Dialog

Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols

2026-04-20 · Fernando Reitich arxiv

Large language models are increasingly deployed as protocols: structured multi-call procedures that spend additional computation to transform a baseline answer into a final one. These protocols are evaluated only by end-…

The Initial Screening Order Problem

2023-07-28 · Jose M. Alvarez, Antonio Mastropietro, Salvatore Ruggieri

We investigate the role of the initial screening order (ISO) in candidate screening. The ISO refers to the order in which the screener searches the candidate pool when selecting $k$ candidates. Today, it is common for th…

Decision MakingFairnessPosition

Gaussian-Bernoulli RBMs Without Tears

2022-10-19 · Renjie Liao, Simon Kornblith, Mengye Ren, David J. Fleet 외

We revisit the challenging problem of training Gaussian-Bernoulli restricted Boltzmann machines (GRBMs), introducing two innovations. We propose a novel Gibbs-Langevin sampling algorithm that outperforms existing methods…