paper-with-me

홈 › Papers

Counteracts: Testing Stereotypical Representation in Pre-trained Language Models

2023-01-11 · Damin Zhang, Julia Rayz, Romila Pradhan

Recently, language models have demonstrated strong performance on various natural language understanding tasks. Language models trained on large human-generated corpus encode not only a significant amount of human knowledge, but also the human stereotype. As more and more downstream tasks have integrated language models as part of the pipeline, it is necessary to understand the internal stereotypical representation in order to design the methods for mitigating the negative effects. In this paper, we use counterexamples to examine the internal stereotypical knowledge in pre-trained language models (PLMs) that can lead to stereotypical preference. We mainly focus on gender stereotypes, but the method can be extended to other types of stereotype. We evaluate 7 PLMs on 9 types of cloze-style prompt with different information and base knowledge. The results indicate that PLMs show a certain amount of robustness against unrelated information and preference of shallow linguistic cues, such as word position and syntactic structure, but a lack of interpreting information by meaning. Such findings shed light on how to interact with PLMs in a neutral approach for both finetuning and evaluation.

📄 PDF Abstract BibTeX arXiv:2301.04347

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Understanding

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Characterizing Stereotypical Bias from Privacy-preserving Pre-Training

2024-06-30 · Stefan Arnold, Rene Gröbner, Annika Schreiner

Differential Privacy (DP) can be applied to raw text by exploiting the spatial arrangement of words in an embedding space. We investigate the implications of such text privatization on Language Models (LMs) and their ten…

Language ModelingLanguage ModellingPrivacy Preserving

Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence Reasoning

2023-03-10 · Hongyin Luo, James Glass

Due to their similarity-based learning objectives, pretrained sentence encoders often internalize stereotypical assumptions that reflect the social biases that exist within their training corpora. In this paper, we descr…

Natural Language InferenceSentencetext similarity

StereoSet: Measuring stereotypical bias in pretrained language models

2020-04-20 · ACL 2021 5 · Moin Nadeem, Anna Bethke, Siva Reddy

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language mod…

Bias DetectionMath

Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models

2023-09-15 · Khyati Khandelwal, Manuel Tonneau, Andrew M. Bean, Hannah Rose Kirk 외

Large Language Models (LLMs), now used daily by millions, can encode societal biases, exposing their users to representational harms. A large body of scholarship on LLM bias exists but it predominantly adopts a Western-c…

FairnessLanguage ModellingLarge Language Model

Testing RadiX-Nets: Advances in Viable Sparse Topologies

2023-11-06 · Kevin Kwak, Zack West, Hayden Jananthan, Jeremy Kepner

The exponential growth of data has sparked computational demands on ML research and industry use. Sparsification of hyper-parametrized deep neural networks (DNNs) creates simpler representations of complex data. Past res…