paper-with-me

홈 › Papers

MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors

2025-09-04 · Xin Tong, Zhi Lin, Jingya Wang, Meng Han, Bo Jin arxiv

Large language models (LLMs) enforce safety alignment to reliably refuse malicious requests, yet the same blanket safeguards also block legitimate uses in policing, defense, and other high-stakes settings. Earlier "refusal-direction" edits can bypass those layers, but they rely on a single vector that indiscriminately unlocks all hazardous topics, offering no semantic control. We introduce Mutually Exclusive Unlock Vectors (MEUV), a lightweight framework that factorizes the monolithic refusal direction into topic-aligned, nearly orthogonal vectors, each dedicated to one sensitive capability. MEUV is learned in a single epoch with a multi-task objective that blends a differential-ablation margin, cross-topic and orthogonality penalties, and several auxiliary terms. On bilingual malicious-prompt benchmarks, MEUV achieves an attack success rate of no less than 87% on Gemma-2-2B, LLaMA-3-8B, and Qwen-7B, yet cuts cross-topic leakage by up to 90% compared with the best single-direction baseline. Vectors trained in Chinese transfer almost unchanged to English (and vice versa), suggesting a language-agnostic refusal subspace. The results show that fine-grained, topic-level capability activation is achievable with minimal utility loss, paving the way for controlled LLMs deployment in security-sensitive domains.

📄 PDF Abstract BibTeX arXiv:2509.12221

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation

2025-10-25 · Ling Team, Ang Li, Ben Liu, Binbin Hu 외 arxiv

We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of billions to one trillion parameters under …

Computational Efficiency

Dual Grained Quantization: Efficient Fine-Grained Quantization for LLM

2023-10-07 · Luoming Zhang, Wen Fei, Weijia Wu, Yefei He 외

Large Language Models (LLMs) pose significant hardware challenges related to memory requirements and computational ability. There are two mainstream quantization schemes for LLMs: coarse-grained ($\textit{e.g.,}$ channel…

Quantization

Fine-Grained Activation Steering: Steering Less, Achieving More

2026-02-04 · Zijian Feng, Tianjiao Li, Zixiao Zhu, Hanzhang Zhou 외 arxiv

Activation steering has emerged as a cost-effective paradigm for modifying large language model (LLM) behaviors. Existing methods typically intervene at the block level, steering the bundled activations of selected atten…

MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training

2024-07-16 · Pinxue Zhao, Hailin Zhang, Fangcheng Fu, Xiaonan Nie 외

Nowadays, Large Language Models (LLMs) have been trained using extended context lengths to foster more creative applications. However, long context training poses great challenges considering the constraint of GPU memory…

CPUGPUManagement

Achieving binary weight and activation for LLMs using Post-Training Quantization

2025-04-07 · Siqing Song, Chuang Wang, Ruiqi Wang, Yi Yang 외

Quantizing large language models (LLMs) to 1-bit precision significantly reduces computational costs, but existing quantization techniques suffer from noticeable performance degradation when using weight and activation p…

Quantization