Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
Recent advancements in LLMs have raised significant safety concerns, particularly when dealing with code-mixed inputs and outputs. Our study systematically investigates the increased susceptibility of LLMs to produce unsafe outputs from code-mixed prompts compared to monolingual English prompts. Utilizing explainability methods, we dissect the internal attribution shifts causing model's harmful behaviors. In addition, we explore cultural dimensions by distinguishing between universally unsafe and culturally-specific unsafe queries. This paper presents novel experimental insights, clarifying the mechanisms driving this phenomenon.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs
Large language models (LLMs) are increasingly examined as both behavioral subjects and decision systems, yet it remains unclear whether observed cognitive biases reflect surface imitation or deeper probability shifts. An…
Inferring Latent Intentions: Attributional Natural Language Inference in LLM Agents
Attributional inference, the ability to predict latent intentions behind observed actions, is a critical yet underexplored capability for large language models (LLMs) operating in multi-agent environments. Traditional na…
Natural Language InferenceFAR: A General Framework for Attributional Robustness
Attribution maps are popular tools for explaining neural networks predictions. By assigning an importance value to each input dimension that represents its impact towards the outcome, they give an intuitive explanation o…
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversatio…
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
Safety tuning through supervised fine-tuning and reinforcement learning from human feedback has substantially improved the robustness of large language models (LLMs). However, it often suppresses rather than eliminates u…
Reinforcement LearningResponse Generation