paper-with-me

Papers

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

2025-09-11 · Jiaqi Weng, Han Zheng, Hanyu Zhang, Ej Zhou, Qinqin He, Jialing Tao, Hui Xue, Zhixuan Chu, Xiting Wang arxiv

Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, under what circumstances SAEs derive most fine-grained latent features for safety, a low-frequency concept domain, remains unexplored. Two key challenges exist: identifying SAEs with the greatest potential for generating safety domain-specific features, and the prohibitively high cost of detailed feature explanation. In this paper, we propose Safe-SAIL, a unified framework for interpreting SAE features in safety-critical domains to advance mechanistic understanding of large language models. Safe-SAIL introduces a pre-explanation evaluation metric to efficiently identify SAEs with strong safety domain-specific interpretability, and reduces interpretation cost by 55% through a segment-level simulation strategy. Building on Safe-SAIL, we train a comprehensive suite of SAEs with human-readable explanations and systematic evaluations for 1,758 safety-related features spanning four domains: pornography, politics, violence, and terror. Using this resource, we conduct empirical analyses and provide insights on the effectiveness of Safe-SAIL for risk feature identification and how safety-critical entities and concepts are encoded across model layers. All models, explanations, and tools are publicly released in our open-source toolkit and companion product.

📄 PDF Abstract BibTeX arXiv:2509.18127

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation

2025-12-29 · Dianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu 외 arxiv

Safety alignment -- training large language models (LLMs) to refuse harmful requests while remaining helpful -- is critical for responsible deployment. Prior work established that safety behaviors are governed by low-ran…

parameter-efficient fine-tuningReinforcement Learning

Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models

2024-05-27 · Shengyun Peng, Pin-Yu Chen, Matthew Hull, Duen Horng Chau

Safety alignment is crucial to ensure that large language models (LLMs) behave in ways that align with human preferences and prevent harmful actions during inference. However, recent studies show that the alignment can b…

Safety Alignment

Safe Reinforcement Learning Using Advantage-Based Intervention

2021-06-16 · Nolan Wagener, Byron Boots, Ching-An Cheng

Many sequential decision problems involve finding a policy that maximizes total reward while obeying safety constraints. Although much recent research has focused on the development of safe reinforcement learning (RL) al…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement Learning

Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models

2025-03-24 · Pin-Yu Chen, Han Shen, Payel Das, Tianyi Chen

Fine-tuning Large Language Models (LLMs) on some task-specific datasets has been a primary use of LLMs. However, it has been empirically observed that this approach to enhancing capability inevitably compromises safety, …

Vision Language Model for Interpretable and Fine-grained Detection of Safety Compliance in Diverse Workplaces

2024-08-13 · Zhiling Chen, Hanning Chen, Mohsen Imani, Ruimin Chen 외

Workplace accidents due to personal protective equipment (PPE) non-compliance raise serious safety concerns and lead to legal liabilities, financial penalties, and reputational damage. While object detection models have …

AttributeLanguage ModelingLanguage Modellingobject-detection+3