paper-with-me

Papers

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

2025-09-18 · Yujia Hu, Ming Shan Hee, Preslav Nakov, Roy Ka-Wei Lee arxiv

The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap. In particular, we introduce \textsf{SGToxicGuard}, a novel dataset and evaluation framework for benchmarking LLM safety in Singapore's diverse linguistic context, including Singlish, Chinese, Malay, and Tamil. SGToxicGuard adopts a red-teaming approach to systematically probe LLM vulnerabilities in three real-world scenarios: \textit{conversation}, \textit{question-answering}, and \textit{content composition}. We conduct extensive experiments with state-of-the-art multilingual LLMs, and the results uncover critical gaps in their safety guardrails. By offering actionable insights into cultural sensitivity and toxicity mitigation, we lay the foundation for safer and more inclusive AI systems in linguistically diverse environments.\footnote{Link to the dataset: https://github.com/Social-AI-Studio/SGToxicGuard.} \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}

📄 PDF Abstract BibTeX arXiv:2509.15260

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RabakBench: Scaling Human Annotations to Construct Localized Multilingual Safety Benchmarks for Low-Resource Languages

2025-07-08 · Gabriel Chua, Leanne Tan, Ziyu Ge, Roy Ka-Wei Lee

Large language models (LLMs) and their safety classifiers often perform poorly on low-resource languages due to limited training data and evaluation benchmarks. This paper introduces RabakBench, a new multilingual safety…

Red Teaming

Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation

2025-07-16 · Ziyu Ge, Gabriel Chua, Leanne Tan, Roy Ka-Wei Lee arxiv

As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing, and culturally embedded markers of harm…

Semantic SimilarityPrompt Engineering

ASTPrompter: Weakly Supervised Automated Language Model Red-Teaming to Identify Low-Perplexity Toxic Prompts

2024-07-12 · Amelia F. Hardy, Houjun Liu, Bernard Lange, Duncan Eddy 외

Conventional approaches for the automated red-teaming of large language models (LLMs) aim to identify prompts that elicit toxic outputs from a frozen language model (the defender). This often results in the prompting mod…

Language ModelingLanguage ModellingRed Teaming

PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models

2024-05-15 · Devansh Jain, Priyanshu Kumar, Samuel Gehman, Xuhui Zhou 외

Recent advances in large language models (LLMs) have led to their extensive global deployment, and ensuring their safety calls for comprehensive and multilingual toxicity evaluations. However, existing toxicity benchmark…

Benchmarking

Learning diverse attacks on large language models for robust red-teaming and safety tuning

2024-05-28 · Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre 외

Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of large language models (LLMs). Developing effective protection against many modes of…

DiversityLanguage ModelingLanguage ModellingRed Teaming