paper-with-me

홈 › Papers

Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations

2025-05-20 · Somnath Banerjee, Pratyush Chatterjee, Shanu Kumar, Sayan Layek, Parag Agrawal, Rima Hazra, Animesh Mukherjee

Recent advancements in LLMs have raised significant safety concerns, particularly when dealing with code-mixed inputs and outputs. Our study systematically investigates the increased susceptibility of LLMs to produce unsafe outputs from code-mixed prompts compared to monolingual English prompts. Utilizing explainability methods, we dissect the internal attribution shifts causing model's harmful behaviors. In addition, we explore cultural dimensions by distinguishing between universally unsafe and culturally-specific unsafe queries. This paper presents novel experimental insights, clarifying the mechanisms driving this phenomenon.

📄 PDF Abstract BibTeX arXiv:2505.14469

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs

2025-11-07 · Felipe Valencia-Clavijo arxiv

Large language models (LLMs) are increasingly examined as both behavioral subjects and decision systems, yet it remains unclear whether observed cognitive biases reflect surface imitation or deeper probability shifts. An…

Inferring Latent Intentions: Attributional Natural Language Inference in LLM Agents

2026-01-13 · Xin Quan, Jiafeng Xiong, Marco Valentino, André Freitas arxiv

Attributional inference, the ability to predict latent intentions behind observed actions, is a critical yet underexplored capability for large language models (LLMs) operating in multi-agent environments. Traditional na…

Natural Language Inference

FAR: A General Framework for Attributional Robustness

2020-10-14 · Adam Ivankay, Ivan Girardi, Chiara Marchiori, Pascal Frossard

Attribution maps are popular tools for explaining neural networks predictions. By assigning an importance value to each input dimension that represents its impact towards the outcome, they give an intuitive explanation o…

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

2026-03-15 · Pengcheng Li, Jie Zhang, Tianwei Zhang, Han Qiu 외 arxiv

Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversatio…

Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling

2026-03-15 · Suvadeep Hajra, Palash Nandi, Tanmoy Chakraborty arxiv

Safety tuning through supervised fine-tuning and reinforcement learning from human feedback has substantially improved the robustness of large language models (LLMs). However, it often suppresses rather than eliminates u…

Reinforcement LearningResponse Generation