paper-with-me

홈 › Papers

Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

2024-01-20 · Kuofeng Gao, Yang Bai, Jindong Gu, Shu-Tao Xia, Philip Torr, Zhifeng Li, Wei Liu

Large vision-language models (VLMs) such as GPT-4 have achieved exceptional performance across various multi-modal tasks. However, the deployment of VLMs necessitates substantial energy consumption and computational resources. Once attackers maliciously induce high energy consumption and latency time (energy-latency cost) during inference of VLMs, it will exhaust computational resources. In this paper, we explore this attack surface about availability of VLMs and aim to induce high energy-latency cost during inference of VLMs. We find that high energy-latency cost during inference of VLMs can be manipulated by maximizing the length of generated sequences. To this end, we propose verbose images, with the goal of crafting an imperceptible perturbation to induce VLMs to generate long sentences during inference. Concretely, we design three loss objectives. First, a loss is proposed to delay the occurrence of end-of-sequence (EOS) token, where EOS token is a signal for VLMs to stop generating further tokens. Moreover, an uncertainty loss and a token diversity loss are proposed to increase the uncertainty over each generated token and the diversity among all tokens of the whole generated sequence, respectively, which can break output dependency at token-level and sequence-level. Furthermore, a temporal weight adjustment algorithm is proposed, which can effectively balance these losses. Extensive experiments demonstrate that our verbose images can increase the length of generated sequences by 7.87 times and 8.56 times compared to original images on MS-COCO and ImageNet datasets, which presents potential challenges for various applications. Our code is available at https://github.com/KuofengGao/Verbose_Images.

📄 PDF Abstract BibTeX arXiv:2401.11170

Code (1)

kuofenggao/verbose_images 공식 구현 pytorch

Tasks

Diversity

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.

Similar Papers 제목 키워드 기반

LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation

2025-11-11 · Xingyu Li, Xiaolei Liu, Cheng Liu, Yixiao Xu 외 arxiv

As large language models (LLMs) scale, their inference incurs substantial computational resources, exposing them to energy-latency attacks, where crafted prompts induce high energy and latency cost. Existing attack metho…

ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models

2025-10-31 · Xin Tang, Youfang Han, Fangfei Gou, Wei Zhao 외 arxiv

Vision-Language Models (VLMs) excel in diverse multimodal tasks. However, user requirements vary across scenarios, which can be categorized into fast response, high-quality output, and low energy consumption. Relying sol…

Sponge Examples: Energy-Latency Attacks on Neural Networks

2020-06-05 · Ilia Shumailov, Yiren Zhao, Daniel Bates, Nicolas Papernot 외

The high energy costs of neural network training and inference led to the use of acceleration hardware such as GPUs and TPUs. While this enabled us to train large-scale neural networks in datacenters and deploy them on e…

Autonomous Vehicles

Energy-Efficient Visual Search by Eye Movement and Low-Latency Spiking Neural Network

2023-10-10 · Yunhui Zhou, Dongqi Han, Yuguo Yu

Human vision incorporates non-uniform resolution retina, efficient eye movement strategy, and spiking neural network (SNN) to balance the requirements in visual field size, visual resolution, energy cost, and inference l…

Green Time-Critical Fog Communication and Computing

2023-03-23 · Hanna Bogucka, Bartosz Kopras, Filip Idzikowski, Bartosz Bossy 외

Fog computing allows computationally-heavy problems with tight time constraints to be solved even if end devices have limited computational resources and latency induced by cloud computing is too high. How can energy con…

Cloud Computing