paper-with-me

Papers

SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models

2025-09-26 · Jingkai Guo, Chaitali Chakrabarti, Deliang Fan arxiv

Model integrity of Large language models (LLMs) has become a pressing security concern with their massive online deployment. Prior Bit-Flip Attacks (BFAs) -- a class of popular AI weight memory fault-injection techniques -- can severely compromise Deep Neural Networks (DNNs): as few as tens of bit flips can degrade accuracy toward random guessing. Recent studies extend BFAs to LLMs and reveal that, despite the intuition of better robustness from modularity and redundancy, only a handful of adversarial bit flips can also cause LLMs' catastrophic accuracy degradation. However, existing BFA methods typically focus on either integer or floating-point models separately, limiting attack flexibility. Moreover, in floating-point models, random bit flips often cause perturbed parameters to extreme values (e.g., flipping in exponent bit), making it not stealthy and leading to numerical runtime error (e.g., invalid tensor values (NaN/Inf)). In this work, for the first time, we propose SBFA (Sneaky Bit-Flip Attack), which collapses LLM performance with only one single bit flip while keeping perturbed values within benign layer-wise weight distribution. It is achieved through iterative searching and ranking through our defined parameter sensitivity metric, ImpactScore, which combines gradient sensitivity and perturbation range constrained by the benign layer-wise weight distribution. A novel lightweight SKIP searching algorithm is also proposed to greatly reduce searching complexity, which leads to successful SBFA searching taking only tens of minutes for SOTA LLMs. Across Qwen, LLaMA, and Gemma models, with only one single bit flip, SBFA successfully degrades accuracy to below random levels on MMLU and SST-2 in both BF16 and INT8 data formats. Remarkably, flipping a single bit out of billions of parameters reveals a severe security concern of SOTA LLM models.

📄 PDF Abstract BibTeX arXiv:2509.21843

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SneakyPrompt: Jailbreaking Text-to-image Generative Models

2023-05-20 · Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong 외

Text-to-image generative models such as Stable Diffusion and DALL$\cdot$E raise many ethical concerns due to the generation of harmful images such as Not-Safe-for-Work (NSFW) ones. To address these ethical concerns, safe…

Reinforcement Learning (RL)Semantic SimilaritySemantic Textual Similarity

NeuroAttack: Undermining Spiking Neural Networks Security through Externally Triggered Bit-Flips

2020-05-16 · Valerio Venceslai, Alberto Marchisio, Ihsen Alouani, Maurizio Martina 외

Due to their proven efficiency, machine-learning systems are deployed in a wide range of complex real-life problems. More specifically, Spiking Neural Networks (SNNs) emerged as a promising solution to the accuracy, reso…

BIG-bench Machine Learning

FlipAttack: Jailbreak LLMs via Flipping

2024-10-02 · Yue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu 외

This paper proposes a simple yet effective jailbreak attack named FlipAttack against black-box LLMs. First, from the autoregressive nature, we reveal that LLMs tend to understand the text from left to right and find that…

PrisonBreak: Jailbreaking Large Language Models with Fewer Than Twenty-Five Targeted Bit-flips

2024-12-10 · Zachary Coalson, Jeonghyun Woo, Yu Sun, Shiyang Chen 외

We introduce a new class of attacks on commercial-scale (human-aligned) language models that induce jailbreaking through targeted bitwise corruptions in model parameters. Our adversary can jailbreak billion-parameter lan…

Computational Efficiency

Alphabet Index Mapping: Jailbreaking LLMs through Semantic Dissimilarity

2025-06-15 · Bilal Saleh Husain

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their susceptibility to adversarial attacks, particularly jailbreaking, poses significant safety and ethical concerns. While numerous jailbreak …

Adversarial Attack