paper-with-me

홈 › Papers

Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models

2026-07-02 · Zikai Zhang, Rui Hu, Olivera Kotevska, Jiahao Xu arxiv

Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and cloud servers. In this process, intermediate vision tokens are transmitted from the edge to the cloud over a communication link, thereby exposing a new attack surface. We study vision token manipulation attack (VTM-Attack) under a black-box man-in-the-middle setting, where an adversary intercepts and manipulates a subset of transmitted vision tokens under a budget constraint. We propose four naïve attack strategies and an optimization-based token selection method. Experiments on 6 state-of-the-art LVLMs (3B-72B) across 4 benchmarks show that manipulating only 10\% of vision tokens can reduce accuracy by up to 88.31\%. These results reveal a critical vulnerability in cloud-edge LVLM inference.

📄 PDF Abstract BibTeX arXiv:2607.02819

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality

2025-01-22 · Yanming Xiu, Tim Scargill, Maria Gorlatova

In Augmented Reality (AR), virtual content enhances user experience by providing additional information. However, improperly positioned or designed virtual content can be detrimental to task performance, as it can impair…

Language ModelingLanguage Modelling

Simple Transparent Adversarial Examples

2021-05-20 · Jaydeep Borkar, Pin-Yu Chen

There has been a rise in the use of Machine Learning as a Service (MLaaS) Vision APIs as they offer multiple services including pre-built models and algorithms, which otherwise take a huge amount of resources if built fr…

Image Generationobject-detectionObject DetectionOptical Character Recognition+1

Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models

2026-05-17 · Zhiqiang Wang, Dongrui Liu, Yan Li, Zonghao Ying 외 arxiv

Existing adversarial attacks on vision-language models (VLMs) can steer model outputs toward attacker-specified target responses, but their effectiveness often degrades when the same perturbed input is paired with differ…

Response GenerationAdversarial Attack

LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference

2026-07-30 · Feng Yang, Xinrui Ju, Keyang Zhang, Xiandong Meng 외 arxiv

Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge devices encode visual inputs into tokens for a general-purpose cloud …

Federated Learning for Large-Scale Cloud Robotic Manipulation: Opportunities and Challenges

2025-07-23 · Obaidullah Zaland, Chanh Nguyen, Florian T. Pokorny, Monowar Bhuyan arxiv

Federated Learning (FL) is an emerging distributed machine learning paradigm, where the collaborative training of a model involves dynamic participation of devices to achieve broad objectives. In contrast, classical mach…

Federated Learning