paper-with-me

Papers

Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models

2023-09-08 · Arka Dutta, Adel Khorramrouz, Sujan Dutta, Ashiqur R. KhudaBukhsh

This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide suite of large language models. Spanning a set of 1,266 identity groups, we first conduct a bias audit of \texttt{PaLM 2} guardrails presenting key insights. Next, we report generalizability across several other models. Through the elicited toxic content, we present a broad analysis with a key emphasis on racism, antisemitism, misogyny, Islamophobia, homophobia, and transphobia. Finally, driven by concrete examples, we discuss potential ramifications.

📄 PDF Abstract BibTeX arXiv:2309.06415

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

PaLM 설명 없음

Similar Papers 제목 키워드 기반

On Bias and Fairness in NLP: Investigating the Impact of Bias and Debiasing in Language Models on the Fairness of Toxicity Detection

2023-05-22 · Fatma Elsafoury, Stamos Katsigiannis

Language models are the new state-of-the-art natural language processing (NLP) models and they are being increasingly used in many NLP tasks. Even though there is evidence that language models are biased, the impact of t…

ClassificationFairnessSelection biastext-classification+1

Black Holes and White Rabbits: Metaphor Identification with Visual Features

2016-06-01 · NAACL 2016 6 · Ekaterina Shutova, Douwe Kiela, Jean Maillard
Clustering

Classifying active and inactive states of growing rabbits from accelerometer data using machine learning algorithms

2024-06-28 · Mónica Mora, Lucile Riaboff, Ingrid David, Juan Pablo Sánchez 외

This study explores how wearable accelerometers, small devices that measure acceleration, can help monitor the activity of growing rabbits. We equipped 16 rabbits with these devices and filmed them for two weeks. By watc…

Management

LED down the rabbit hole: exploring the potential of global attention for biomedical multi-document summarisation

2022-09-19 · sdp (COLING) 2022 10 · Yulia Otmakhova, Hung Thinh Truong, Timothy Baldwin, Trevor Cohn 외

In this paper we report on our submission to the Multidocument Summarisation for Literature Review (MSLR) shared task. Specifically, we adapt PRIMERA (Xiao et al., 2022) to the biomedical domain by placing global attenti…

What-if I ask you to explain: Explaining the effects of perturbations in procedural text

2020-05-04 · Findings of the Association for Computational Linguistics 2020 · Dheeraj Rajagopal, Niket Tandon, Bhavana Dalvi, Peter Clark 외

We address the task of explaining the effects of perturbations in procedural text, an important test of process comprehension. Consider a passage describing a rabbit's life-cycle: humans can easily explain the effect on …