paper-with-me

홈 › Papers

SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification

2025-06-20 · ZhengLin Lai, MengYao Liao, Dong Xu, Zebin Zhao, Zhihang Yuan, Chao Fan, Jianqiang Li, Bingzhe Wu

Large language models based on Mixture-of-Experts have achieved substantial gains in efficiency and scalability, yet their architectural uniqueness introduces underexplored safety alignment challenges. Existing safety alignment strategies, predominantly designed for dense models, are ill-suited to address MoE-specific vulnerabilities. In this work, we formalize and systematically study MoE model's positional vulnerability - the phenomenon where safety-aligned behaviors rely on specific expert modules, revealing critical risks inherent to MoE architectures. To this end, we present SAFEx, an analytical framework that robustly identifies, characterizes, and validates the safety-critical experts using a novel Stability-based Expert Selection (SES) algorithm. Notably, our approach enables the explicit decomposition of safety-critical experts into distinct functional groups, including those responsible for harmful content detection and those controlling safe response generation. Extensive experiments on mainstream MoE models, such as the recently released Qwen3-MoE, demonstrated that their intrinsic safety mechanisms heavily rely on a small subset of positional experts. Disabling these experts significantly compromised the models' ability to refuse harmful requests. For Qwen3-MoE with 6144 experts (in the FNN layer), we find that disabling as few as 12 identified safety-critical experts can cause the refusal rate to drop by 22%, demonstrating the disproportionate impact of a small set of experts on overall model safety.

📄 PDF Abstract BibTeX arXiv:2506.17368

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-ExpertsResponse GenerationSafety Alignment

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
MoE 설명 없음

Similar Papers 제목 키워드 기반

LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models

2024-07-23 · Shi Lin, Hongming Yang, Rongchang Li, Xun Wang 외

The rapid development of Large Language Models (LLMs) has brought impressive advancements across various tasks. However, despite these achievements, LLMs still pose inherent safety risks, especially in the context of jai…

Multimodal ReasoningPrompt EngineeringVisual Reasoning

Explainable AI for Safe and Trustworthy Autonomous Driving: A Systematic Review

2024-02-08 · Anton Kuznietsov, Balint Gyevnar, Cheng Wang, Steven Peters 외

Artificial Intelligence (AI) shows promising applications for the perception and planning tasks in autonomous driving (AD) due to its superior performance compared to conventional methods. However, inscrutable AI systems…

Autonomous DrivingSystematic Literature Review

Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks

2025-02-18 · Wenpeng Xing, Minghao Li, Mohan Li, Meng Han

Embodied AI systems, including robots and autonomous vehicles, are increasingly integrated into real-world applications, where they encounter a range of vulnerabilities stemming from both environmental and system-level f…

Adversarial AttackAutonomous VehiclesDecision MakingMotion Planning+2

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

2026-01-07 · Xinyue Lou, Jinan Xu, Jingyi Yin, Xiaolong Wang 외 arxiv

As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to human behavior, perpetually overhanging human society like a sword of D…

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

2025-09-04 · Chuhan Zhang, Ye Zhang, Bowen Shi, Yuyou Gan 외 arxiv

In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to byp…