Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide suite of large language models. Spanning a set of 1,266 identity groups, we first conduct a bias audit of \texttt{PaLM 2} guardrails presenting key insights. Next, we report generalizability across several other models. Through the elicited toxic content, we present a broad analysis with a key emphasis on racism, antisemitism, misogyny, Islamophobia, homophobia, and transphobia. Finally, driven by concrete examples, we discuss potential ramifications.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
On Bias and Fairness in NLP: Investigating the Impact of Bias and Debiasing in Language Models on the Fairness of Toxicity Detection
Language models are the new state-of-the-art natural language processing (NLP) models and they are being increasingly used in many NLP tasks. Even though there is evidence that language models are biased, the impact of t…
ClassificationFairnessSelection biastext-classification+1Black Holes and White Rabbits: Metaphor Identification with Visual Features
Classifying active and inactive states of growing rabbits from accelerometer data using machine learning algorithms
This study explores how wearable accelerometers, small devices that measure acceleration, can help monitor the activity of growing rabbits. We equipped 16 rabbits with these devices and filmed them for two weeks. By watc…
ManagementLED down the rabbit hole: exploring the potential of global attention for biomedical multi-document summarisation
In this paper we report on our submission to the Multidocument Summarisation for Literature Review (MSLR) shared task. Specifically, we adapt PRIMERA (Xiao et al., 2022) to the biomedical domain by placing global attenti…
What-if I ask you to explain: Explaining the effects of perturbations in procedural text
We address the task of explaining the effects of perturbations in procedural text, an important test of process comprehension. Consider a passage describing a rabbit's life-cycle: humans can easily explain the effect on …