paper-with-me

홈 › Papers

Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining

2025-08-08 · Bing Han, Feifei Zhao, Dongcheng Zhao, Guobin Shen, Ping Wu, Yu Shi, Yi Zeng arxiv

While fine-tuning services drive the rapid expansion of task capabilities in large language models (LLMs), they are often accompanied by the degradation and reorganization of safety-aligned representations, making models more prone to deviating from human preferences and exposing them to emerging jailbreak risks. Existing post-fine-tuning defense methods predominantly rely on single-scale safety correction mechanisms, which struggle to achieve a robust balance among safety, model utility, and continual adaptability. We propose Multi-Level Safety Continual Projection (MSCP), a training-free post-fine-tuning safety enhancement method that implicitly aligns global and localized safety activations through coordinated multi-level representations to isolate sparse neuron clusters governing safety-sensitive behaviors. It then applies composable safety-direction projections without retraining, effectively suppressing harmful outputs under minimal parameter perturbations while preserving task performance and improving alignment with human preferences. Extensive experiments across multiple fine-tuned LLM models demonstrate that our method significantly reduce harmfulness scores and attack success rates with minimal parameter modifications, while preserving the model's utility. Furthermore, we introduce a task-specific, multi-dimensional heterogeneous safety activation clustering mechanism that enables continual defense and generalization capability against unforeseen emerging safety concerns.

📄 PDF Abstract BibTeX arXiv:2508.09190

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning

2026-05-13 · Timofey Tomashevskiy arxiv

Safety in reinforcement learning is often specified through cumulative cost constraints, but these trajectory-level guarantees do not directly prevent unsafe individual decisions, especially under nonstationarity. In con…

Reinforcement Learning

LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation

2025-04-10 · Juzheng Zhang, Jiacheng You, Ashwinee Panda, Tom Goldstein

Low-Rank Adaptation (LoRA) has emerged as a popular parameter-efficient fine-tuning (PEFT) method for Large Language Models (LLMs), yet it still incurs notable overhead and suffers from parameter interference in multi-ta…

Code GenerationContinual LearningMathematical ReasoningNatural Language Understanding+2

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning

2026-05-06 · Xiao Wang, Yifei Zhang, YongKang Liu, Xiaocui Yang 외 arxiv

Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors learned from millions of preference examples. Existing studies attempt …

Continual Gradient Low-Rank Projection Fine-Tuning for LLMs

2025-07-03 · Chenxu Wang, Yilin Lyu, Zicheng Sun, Liping Jing

Continual fine-tuning of Large Language Models (LLMs) is hampered by the trade-off between efficiency and expressiveness. Low-Rank Adaptation (LoRA) offers efficiency but constrains the model's ability to learn new tasks…

Continual Learning

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

2026-02-08 · Guanglong Sun, Siyuan Zhang, Liyuan Wang, Jun Zhu 외 arxiv

Safety post-training can improve the harmfulness and policy compliance of Large Language Models (LLMs), but it may also reduce general utility, a phenomenon often described as the \emph{alignment tax}. We study this trad…

Continual Learning