Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs
Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classification proves insufficient when models encounter ethical dilemmas, where the capacity to reason through moral trade-offs creates a distinct attack surface. We formalize this vulnerability through TRIAL, a multi-turn red-teaming methodology that embeds harmful requests within ethical framings. TRIAL achieves high attack success rates across most tested models by systematically exploiting the model's ethical reasoning capabilities to frame harmful actions as morally necessary compromises. Building on these insights, we introduce ERR (Ethical Reasoning Robustness), a defense framework that distinguishes between instrumental responses that enable harmful outcomes and explanatory responses that analyze ethical frameworks without endorsing harmful acts. ERR employs a Layer-Stratified Harm-Gated LoRA architecture, achieving robust defense against reasoning-based attacks while preserving model utility.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
PERC: Posit Enhanced Rocket Chip
Balancing between precision and performance is a known trade-off in system design. Universal Number system attempts to dissolve this trade-off with its flexible arbitrary precision bound within fixed bit length. Its Type…
A Development Cycle for Automated Self-Exploration of Robot Behaviors
In this paper we introduce Q-Rock, a development cycle for the automated self-exploration and qualification of robot behaviors. With Q-Rock, we suggest a novel, integrative approach to automate robot development processe…
BIG-bench Machine LearningMulti-Class Electrical and Mechanical Fault Classification Using Random Convolutional Kernels
Diagnosing faults in rotating machinery is essential for ensuring the reliability of industrial processes. Random convolutional kernel-based Time Series Classification (TSC) methods, such as ROCKET and its variants, prov…
Time Series ClassificationComputational EfficiencyRocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering
In open-domain question answering, dense passage retrieval has become a new paradigm to retrieve relevant passages for finding answers. Typically, the dual-encoder architecture is adopted to learn dense representations o…
Data AugmentationNatural QuestionsOpen-Domain Question AnsweringPassage Retrieval+2Inclusion of Lithological terms (rocks and minerals) in The Open Wordnet for English
We extend the Open WordNet for English (OWN-EN) with rock-related and other lithological terms using the authoritative source of GBA{'}s Thesaurus. Our aim is to improve WordNet to better function within Oil {\&} Gas dom…