paper-with-me

Papers

SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues

2026-03-19 · Carlos Hinojosa, Clemens Grange, Bernard Ghanem arxiv

Vision-language models (VLMs) are increasingly deployed in real-world and embodied settings where safety decisions depend on visual context. However, it remains unclear which visual evidence drives these judgments. We study whether multimodal safety behavior in VLMs can be steered by simple semantic cues. We introduce a semantic steering framework that applies controlled textual, visual, and cognitive interventions without changing the underlying scene content. To evaluate these effects, we propose SAVeS, a benchmark for situational safety under semantic cues, together with an evaluation protocol that separates behavioral refusal, grounded safety reasoning, and false refusals. Experiments across multiple VLMs and an additional state-of-the-art benchmark show that safety decisions are highly sensitive to semantic cues, indicating reliance on learned visual-linguistic associations rather than grounded visual understanding. We further demonstrate that automated steering pipelines can exploit these mechanisms, highlighting a potential vulnerability in multimodal safety systems.

📄 PDF Abstract BibTeX arXiv:2603.19092

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BabelSteering: Multilingual Safety Alignment via English Steering Vectors

2026-08-17 · Emma V. Stein, Dominik Meier, Terry Ruas, Jan Philip Wahle 외 arxiv

Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users interacting with LLMs in other languages may encounte…

Towards Inference-time Category-wise Safety Steering for Large Language Models

2024-10-02 · Amrita Bhattacharjee, Shaona Ghosh, Traian Rebedea, Christopher Parisien

While large language models (LLMs) have seen unprecedented advancements in capabilities and applications across a variety of use-cases, safety alignment of these models is still an area of active research. The fragile na…

Safety Alignment

Evaluating Steering Techniques using Human Similarity Judgments

2025-05-25 · Zach Studdiford, Timothy T. Rogers, Siddharth Suresh, Kushin Mukherjee

Current evaluations of Large Language Model (LLM) steering techniques focus on task-specific performance, overlooking how well steered representations align with human cognition. Using a well-established triadic similari…

Language ModelingLanguage ModellingLarge Language Model

Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models

2026-03-23 · Xingyu Zhu, Beier Zhu, Shuo Wang, Junfeng Fang 외 arxiv

As vision-language models (VLMs) are increasingly deployed in open-world scenarios, they can be easily induced by visual jailbreak attacks to generate harmful content, posing serious risks to model safety and trustworthy…

One-shot Optimized Steering Vector for Hallucination Mitigation for VLMs

2026-01-30 · Youxu Shi, Suorong Yang, Dong Liu arxiv

Vision Language Models (VLMs) achieve strong performance on multimodal tasks but still suffer from hallucination and safety-related failures that persist even at scale. Steering offers a lightweight technique to improve …