paper-with-me

홈 › Papers

Semantics derived automatically from language corpora contain human-like biases

2016-08-25 · Aylin Caliskan, Joanna J. Bryson, Arvind Narayanan

Artificial intelligence and machine learning are in a period of astounding growth. However, there are concerns that these technologies may be used, either with or without intention, to perpetuate the prejudice and unfairness that unfortunately characterizes many human institutions. Here we show for the first time that human-like semantic biases result from the application of standard machine learning to ordinary language---the same sort of language humans are exposed to every day. We replicate a spectrum of standard human biases as exposed by the Implicit Association Test and other well-known psychological studies. We replicate these using a widely used, purely statistical machine-learning model---namely, the GloVe word embedding---trained on a corpus of text from the Web. Our results indicate that language itself contains recoverable and accurate imprints of our historic biases, whether these are morally neutral as towards insects or flowers, problematic as towards race or gender, or even simply veridical, reflecting the {\em status quo} for the distribution of gender with respect to careers or first names. These regularities are captured by machine learning along with the rest of semantics. In addition to our empirical findings concerning language, we also contribute new methods for evaluating bias in text, the Word Embedding Association Test (WEAT) and the Word Embedding Factual Association Test (WEFAT). Our results have implications not only for AI and machine learning, but also for the fields of psychology, sociology, and human ethics, since they raise the possibility that mere exposure to everyday language can account for the biases we replicate here.

📄 PDF Abstract BibTeX arXiv:1608.07187

Code (1)

YunjinPark/modu_project

Tasks

BIG-bench Machine LearningEthicsSociology

Methods 이 논문이 사용한 방법론

GloVe GloVe Embeddings are a type of word embedding that encode the co-occurrence probability ratio between two words as vector differences. GloVe uses a weighted least squares…

Similar Papers 제목 키워드 기반

Software Language Comprehension using a Program-Derived Semantics Graph

2020-04-02 · NeurIPS Workshop CAP 2020 12 · Roshni G. Iyer, Yizhou Sun, Wei Wang, Justin Gottschlich

Traditional code transformation structures, such as abstract syntax trees (ASTs), conteXtual flow graphs (XFGs), and more generally, compiler intermediate representations (IRs), may have limitations in extracting higher-…

Natural Language Semantics With Pictures: Some Language & Vision Datasets and Potential Uses for Computational Semantics

2019-04-15 · David Schlangen

Propelling, and propelled by, the "deep learning revolution", recent years have seen the introduction of ever larger corpora of images annotated with natural language expressions. We survey some of these corpora, taking …

Relation

Natural Language Semantics With Pictures: Some Language \& Vision Datasets and Potential Uses for Computational Semantics

2019-05-01 · WS 2019 5 · David Schlangen

Propelling, and propelled by, the {``}deep learning revolution{''}, recent years have seen the introduction of ever larger corpora of images annotated with natural language expressions. We survey some of these corpora, t…

Relation

Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models

2026-06-15 · Yuanyuan Tian, Wenwen Li, Xiao Chen, Michael Brook 외 arxiv

Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale washed ashore in Alaska) that contain spatial and temporal information, is a critic…

Information RetrievalSemantic Similarity

Enriching a Valency Lexicon by Deverbative Nouns

2016-12-01 · WS 2016 12 · Eva Fu{\v{c}}{\'\i}kov{\'a}, Jan Haji{\v{c}}, Zde{\v{n}}ka Ure{\v{s}}ov{\'a}

We present an attempt to automatically identify Czech deverbative nouns using several methods that use large corpora as well as existing lexical resources. The motivation for the task is to extend a verbal valency (i.e.,…