paper-with-me

홈 › Papers

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

2026-03-09 · Qishun Yang, Shu Yang, Lijie Hu, Di Wang arxiv

Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods require explicit safety labels or contrastive data; yet, threat-related concepts are concrete and visually depictable, while safety concepts, like helpfulness, are abstract and lack visual referents. Inspired by the Self-Fulfilling mechanism underlying emergent misalignment, we propose Visual Self-Fulfilling Alignment (VSFA). VSFA fine-tunes vision-language models (VLMs) on neutral VQA tasks constructed around threat-related images, without any safety labels. Through repeated exposure to threat-related visual content, models internalize the implicit semantics of vigilance and caution, shaping safety-oriented personas. Experiments across multiple VLMs and safety benchmarks demonstrate that VSFA reduces the attack success rate, improves response quality, and mitigates over-refusal while preserving general capabilities. Our work extends the self-fulfilling mechanism from text to visual modalities, offering a label-free approach to VLMs alignment.

📄 PDF Abstract BibTeX arXiv:2603.08486

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap

2025-02-14 · Qin Liu, Fei Wang, Chaowei Xiao, Muhao Chen

The emergence of vision language models (VLMs) comes with increased safety concerns, as the incorporation of multiple modalities heightens vulnerability to attacks. Although VLMs can be built upon LLMs that have textual …

AttributeSafety Alignment

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

2026-01-15 · Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim 외 arxiv

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly…

Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training

2025-10-23 · Zheng-Xin Yong, Stephen H. Bach arxiv

We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaking. Specifically, after benign reasoning training on math or code domains, RLM…

DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models

2025-10-12 · Kaixuan Ren, Preslav Nakov, Usman Naseem arxiv

As vision-language models (VLMs) become increasingly capable, maintaining a balance between safety and usefulness remains a central challenge. Safety mechanisms, while essential, can backfire, causing over-refusal, where…

Shape it Up! Restoring LLM Safety during Finetuning

2025-05-22 · Shengyun Peng, Pin-Yu Chen, Jianfeng Chi, Seongmin Lee 외

Finetuning large language models (LLMs) enables user-specific customization but introduces critical safety risks: even a few harmful examples can compromise safety alignment. A common mitigation strategy is to update the…

Safety Alignment