paper-with-me

홈 › Papers

CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning

2025-05-22 · Biao Yi, Tiansheng Huang, Baolei Zhang, Tong Li, Lihai Nie, Zheli Liu, Li Shen

Fine-tuning-as-a-service, while commercially successful for Large Language Model (LLM) providers, exposes models to harmful fine-tuning attacks. As a widely explored defense paradigm against such attacks, unlearning attempts to remove malicious knowledge from LLMs, thereby essentially preventing them from being used to perform malicious tasks. However, we highlight a critical flaw: the powerful general adaptability of LLMs allows them to easily bypass selective unlearning by rapidly relearning or repurposing their capabilities for harmful tasks. To address this fundamental limitation, we propose a paradigm shift: instead of selective removal, we advocate for inducing model collapse--effectively forcing the model to "unlearn everything"--specifically in response to updates characteristic of malicious adaptation. This collapse directly neutralizes the very general capabilities that attackers exploit, tackling the core issue unaddressed by selective unlearning. We introduce the Collapse Trap (CTRAP) as a practical mechanism to implement this concept conditionally. Embedded during alignment, CTRAP pre-configures the model's reaction to subsequent fine-tuning dynamics. If updates during fine-tuning constitute a persistent attempt to reverse safety alignment, the pre-configured trap triggers a progressive degradation of the model's core language modeling abilities, ultimately rendering it inert and useless for the attacker. Crucially, this collapse mechanism remains dormant during benign fine-tuning, ensuring the model's utility and general capabilities are preserved for legitimate users. Extensive empirical results demonstrate that CTRAP effectively counters harmful fine-tuning risks across various LLMs and attack settings, while maintaining high performance in benign scenarios. Our code is available at https://anonymous.4open.science/r/CTRAP.

📄 PDF Abstract BibTeX arXiv:2505.16559

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelSafety Alignment

Similar Papers 제목 키워드 기반

Towards Building a Robust Toxicity Predictor

2024-04-09 · Dmitriy Bespalov, Sourav Bhabesh, Yi Xiang, Liutong Zhou 외

Recent NLP literature pays little attention to the robustness of toxicity language predictors, while these systems are most likely to be used in adversarial contexts. This paper presents a novel adversarial attack, \text…

Adversarial Attack

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models

2025-01-30 · Haoyu Liang, Youran Sun, Yunfeng Cai, Jun Zhu 외

The security issue of large language models (LLMs) has gained wide attention recently, with various defense mechanisms developed to prevent harmful output, among which safeguards based on text embedding models serve as a…

A Quadratic Speedup in Finding Nash Equilibria of Quantum Zero-Sum Games

2023-11-17 · Francisca Vasconcelos, Emmanouil-Vasileios Vlatakis-Gkaragkounis, Panayotis Mertikopoulos, Georgios Piliouras 외

Recent developments in domains such as non-local games, quantum interactive proofs, and quantum generative adversarial networks have renewed interest in quantum game theory and, specifically, quantum zero-sum games. Cent…

I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift

2026-03-01 · Subramanyam Sahoo, Vinija Jain, Divya Chaudhary, Aman Chadha arxiv

Instruction tuned reasoning models are increasingly deployed with safety classifiers trained on frozen embeddings, assuming representation stability across model updates. We systematically investigate this assumption and…

On the Embedding Collapse when Scaling up Recommendation Models

2023-10-06 · Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen 외

Recent advances in foundation models have led to a promising trend of developing large recommendation models to leverage vast amounts of available data. Still, mainstream models remain embarrassingly small in size and na…

Diversity