paper-with-me

Papers

Fail-Closed Alignment for Large Language Models

2026-02-19 · Zachary Coalson, Beth Sohler, Aiden Gabriel, Sanghyun Hong arxiv

We identify a structural weakness in current large language model (LLM) alignment: modern refusal mechanisms are fail-open. While existing approaches encode refusal behaviors across multiple latent features, suppressing a single dominant feature$-$via prompt-based jailbreaks$-$can cause alignment to collapse, leading to unsafe generation. Motivated by this, we propose fail-closed alignment as a design principle for robust LLM safety: refusal mechanisms should remain effective even under partial failures via redundant, independent causal pathways. We present a concrete instantiation of this principle: a progressive alignment framework that iteratively identifies and ablates previously learned refusal directions, forcing the model to reconstruct safety along new, independent subspaces. Across four jailbreak attacks, we achieve the strongest overall robustness while mitigating over-refusal and preserving generation quality, with small computational overhead. Our mechanistic analyses confirm that models trained with our method encode multiple, causally independent refusal directions that prompt-based jailbreaks cannot suppress simultaneously, providing empirical support for fail-closed alignment as a principled foundation for robust LLM safety.

📄 PDF Abstract BibTeX arXiv:2602.16977

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Tool Manipulation Capability of Open-source Large Language Models

2023-05-25 · Qiantong Xu, Fenglu Hong, Bo Li, Changran Hu 외

Recent studies on software tool manipulation with large language models (LLMs) mostly rely on closed model APIs. The industrial adoption of these models is substantially constrained due to the security and robustness ris…

Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs

2025-02-12 · Mohsinul Kabir, Ajwad Abrar, Sophia Ananiadou

A large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models (LLMs). In this work, we challenge this constrained evaluation paradigm and explore more real…

Multiple-choiceSurvey

ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models

2026-08-21 · Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai, Dawei Feng 외 arxiv

While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging. Multimodal safety alignment methods must address cross-modal jailbreaks, saf…

Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search

2024-08-11 · Robert J. Moss

Eliciting harmful behavior from large language models (LLMs) is an important task to ensure the proper alignment and safety of the models. Often when training LLMs, ethical guidelines are followed yet alignment failures …

Red Teaming

Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling

2026-06-30 · Ziyan Wang, Tan Xiang, Peng Chen, Xintao Yan arxiv

A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in globally observable closed-loop environments. In such logs, the ego vehicle has rich lo…