paper-with-me

홈 › Papers

Sociolectal Analysis of Pretrained Language Models

2021-11-01 · EMNLP 2021 11 · Sheng Zhang, Xin Zhang, Weiming Zhang, Anders Søgaard

Using data from English cloze tests, in which subjects also self-reported their gender, age, education, and race, we examine performance differences of pretrained language models across demographic groups, defined by these (protected) attributes. We demonstrate wide performance gaps across demographic groups and show that pretrained language models systematically disfavor young non-white male speakers; i.e., not only do pretrained language models learn social biases (stereotypical associations) – pretrained language models also learn sociolectal biases, learning to speak more like some than like others. We show, however, that, with the exception of BERT models, larger pretrained language models reduce some the performance gaps between majority and minority groups.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stylistic Evolution and LLM Neutrality in Singlish Language

2026-01-10 · Linus Tze En Foo, Weihan Angela Ng, Wenkai Li, Lynnette Hui Xian Ng arxiv

Singlish is a creole rooted in Singapore's multilingual environment that continues to evolve alongside social and technological change. We examine diachronic stylistic change across a decade of informal digital messages …

Detecting Sexual Content at the Sentence Level in First Millennium Latin Texts

2023-09-25 · Thibault Clérice

In this study, we propose to evaluate the use of deep learning methods for semantic classification at the sentence level to accelerate the process of corpus building in the field of humanities and linguistics, a traditio…

SentenceSentence Classification

Pretrained Transformers as Universal Computation Engines

2021-03-09 · Kevin Lu, Aditya Grover, Pieter Abbeel, Igor Mordatch

We investigate the capability of a transformer pretrained on natural language to generalize to other modalities with minimal finetuning -- in particular, without finetuning of the self-attention and feedforward layers of…

A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language Models

2023-10-19 · Yi Zhou, Jose Camacho-Collados, Danushka Bollegala

Various types of social biases have been reported with pretrained Masked Language Models (MLMs) in prior work. However, multiple underlying factors are associated with an MLM such as its model size, size of the training …

AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages

2022-11-07 · Bonaventure F. P. Dossou, Atnafu Lambebo Tonja, Oreen Yousuf, Salomey Osei 외

In recent years, multilingual pre-trained language models have gained prominence due to their remarkable performance on numerous downstream Natural Language Processing tasks (NLP). However, pre-training these large multi…

Active LearningLanguage ModelingLanguage ModellingNER+3