paper-with-me

홈 › Papers

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

2025-09-07 · Nirmalendu Prakash, Yeo Wei Jie, Amir Abdullah, Ranjan Satapathy, Erik Cambria, Roy Ka Wei Lee arxiv

Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour remain poorly understood. We study two public instruction-tuned models, Gemma-2-2B-IT and LLaMA-3.1-8B-IT, using sparse autoencoders (SAEs) trained on residual-stream activations. Given a harmful prompt, we search the SAE latent space for feature sets whose ablation flips the model from refusal to compliance, demonstrating causal influence and creating a jailbreak. Our search proceeds in three stages: (1) Refusal Direction: find a refusal-mediating direction and collect SAE features near that direction; (2) Greedy Filtering: prune to a minimal set; and (3) Interaction Discovery: fit a factorization machine (FM) that captures nonlinear interactions among the remaining active features and the minimal set. This pipeline yields a broad set of jailbreak-critical features, offering insight into the mechanistic basis of refusal. Moreover, we find evidence of redundant features that remain dormant unless earlier features are suppressed. Our findings highlight the potential for fine-grained auditing and targeted intervention in safety behaviours by manipulating the interpretable latent space.

📄 PDF Abstract BibTeX arXiv:2509.09708

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors

2024-06-20 · Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang 외

Evaluating aligned large language models' (LLMs) ability to recognize and reject unsafe user requests is crucial for safe, policy-compliant deployments. Existing evaluation efforts, however, face three limitations that w…

Language ModelingLanguage ModellingLarge Language Model

FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)

2024-08-28 · Aman Priyanshu, Supriti Vijay

This paper introduces FRACTURED-SORRY-Bench, a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple yet…

The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage

2026-02-10 · Jayesh Choudhari, Piyush Kumar Singh arxiv

Domain fine-tuning is a common path to deploy small instruction-tuned language models as customer-support assistants, yet its effects on safety-aligned behavior and privacy are not well understood. In real deployments, s…

Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering

2026-04-06 · Purva Chiniya, Kevin Scaria, Sagar Chaturvedi arxiv

Large language models (LLMs) remain susceptible to jailbreak and direct prompt-injection attacks, yet the strongest defensive filters frequently over-refuse benign queries and degrade user experience. Previous work on ja…

PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models

2026-06-08 · Gianluca Barmina, Federico Torrielli, Sven Harms, Jacob Nielsen 외 arxiv

Large language models (LLMs) routinely face requests that should be refused, creating a trade-off between helpfulness and harm prevention. However, refusals themselves can be helpful. In high-risk interactions involving …

parameter-efficient fine-tuningDomain Generalization