paper-with-me

홈 › Papers

One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models

2025-07-10 · Jiale Zhao, Xinyang Jiang, Junyao Gao, Yuhao Xue, Cairong Zhao arxiv

Unified vision-language models(VLMs) have recently shown remarkable progress, enabling a single model to flexibly address diverse tasks through different instructions within a shared computational architecture. This instruction-based control mechanism creates unique security challenges, as adversarial inputs must remain effective across multiple task instructions that may be unpredictably applied to process the same malicious content. In this paper, we introduce CrossVLAD, a new benchmark dataset carefully curated from MSCOCO with GPT-4-assisted annotations for systematically evaluating cross-task adversarial attacks on unified VLMs. CrossVLAD centers on the object-change objective-consistently manipulating a target object's classification across four downstream tasks-and proposes a novel success rate metric that measures simultaneous misclassification across all tasks, providing a rigorous evaluation of adversarial transferability. To tackle this challenge, we present CRAFT (Cross-task Region-based Attack Framework with Token-alignment), an efficient region-centric attack method. Extensive experiments on Florence-2 and other popular unified VLMs demonstrate that our method outperforms existing approaches in both overall cross-task attack performance and targeted object-change success rates, highlighting its effectiveness in adversarially influencing unified VLMs across diverse tasks.

📄 PDF Abstract BibTeX arXiv:2507.07709

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Attack

Similar Papers 제목 키워드 기반

Odd-One-Out: Anomaly Detection by Comparing with Neighbors

2024-06-28 · CVPR 2025 1 · Ankan Bhunia, Changjian Li, Hakan Bilen

This paper introduces a novel anomaly detection (AD) problem that focuses on identifying `odd-looking' objects relative to the other instances in a given scene. In contrast to the traditional AD benchmarks, anomalies in …

8kAnomaly DetectionOdd One Out

Scale-Transferrable Object Detection

2018-06-01 · CVPR 2018 6 · Peng Zhou, Bingbing Ni, Cong Geng, Jianguo Hu 외

Scale problem lies in the heart of object detection. In this work, we develop a novel Scale-Transferrable Detection Network (STDN) for detecting multi-scale objects in images. In contrast to previous methods that simply …

Objectobject-detectionObject DetectionSuper-Resolution

Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation

2024-02-02 · Elaf Alhazmi, Quan Z. Sheng, Wei Emma Zhang, Munazza Zaib 외

The distractor generation task focuses on generating incorrect but plausible options for objective questions such as fill-in-the-blank and multiple-choice questions. This task is widely utilized in educational settings a…

Distractor GenerationMultiple-choice

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

2026-07-07 · He Liang, Chenyang Ma, Yiming Zhang, Sangyun Shin 외 arxiv

Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multip…

Graph Neural NetworkScene Understanding

MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects

2024-12-06 · CVPR 2025 1 · Lei Fan, Dongdong Fan, Zhiguang Hu, Yiwen Ding 외

We present MANTA, a visual-text anomaly detection dataset for tiny objects. The visual component comprises over 137.3K images across 38 object categories spanning five typical domains, of which 8.6K images are labeled as…

2kAnomaly DetectionBenchmarkingMultiple-choice