"Im not Racist but...": Discovering Bias in the Internal Knowledge of Large Language Models
Large language models (LLMs) have garnered significant attention for their remarkable performance in a continuously expanding set of natural language processing tasks. However, these models have been shown to harbor inherent societal biases, or stereotypes, which can adversely affect their performance in their many downstream applications. In this paper, we introduce a novel, purely prompt-based approach to uncover hidden stereotypes within any arbitrary LLM. Our approach dynamically generates a knowledge representation of internal stereotypes, enabling the identification of biases encoded within the LLM's internal knowledge. By illuminating the biases present in LLMs and offering a systematic methodology for their analysis, our work contributes to advancing transparency and promoting fairness in natural language processing systems.
Code (0)
등록된 구현이 없습니다.
Tasks
FairnessSimilar Papers 제목 키워드 기반
A Weakly Supervised Classifier and Dataset of White Supremacist Language
We present a dataset and classifier for detecting the language of white supremacist extremism, a growing issue in online hate speech. Our weakly supervised classifier is trained on large datasets of text from explicitly …
Biasly: a machine learning based platform for automatic racial discrimination detection in online texts
Detecting hateful, toxic, and otherwise racist or sexist language in user-generated online contents has become an increasingly important task in recent years. Indeed, the anonymity, transience, size of messages, and the …
BIG-bench Machine LearningManagementA Dictionary-based Approach to Racism Detection in Dutch Social Media
We present a dictionary-based approach to racism detection in Dutch social media comments, which were retrieved from two public Belgian social media sites likely to attract racist reactions. These comments were labeled a…
Machines Do See Color: A Guideline to Classify Different Forms of Racist Discourse in Large Corpora
Current methods to identify and classify racist language in text rely on small-n qualitative approaches or large-n approaches focusing exclusively on overt forms of racist discourse. This article provides a step-by-step …
text-classificationText ClassificationXLM-RDetecting Hate Speech with GPT-3
Sophisticated language models such as OpenAI's GPT-3 can generate hateful text that targets marginalized groups. Given this capacity, we are interested in whether large language models can be used to identify hate speech…
Few-Shot LearningHate Speech DetectionOne-Shot Learning