paper-with-me

홈 › Papers

AraWEAT: Multidimensional Analysis of Biases in Arabic Word Embeddings

2020-11-03 · COLING (WANLP) 2020 12 · Anne Lauscher, Rafik Takieddin, Simone Paolo Ponzetto, Goran Glavaš

Recent work has shown that distributional word vector spaces often encode human biases like sexism or racism. In this work, we conduct an extensive analysis of biases in Arabic word embeddings by applying a range of recently introduced bias tests on a variety of embedding spaces induced from corpora in Arabic. We measure the presence of biases across several dimensions, namely: embedding models (Skip-Gram, CBOW, and FastText) and vector sizes, types of text (encyclopedic text, and news vs. user-generated content), dialects (Egyptian Arabic vs. Modern Standard Arabic), and time (diachronic analyses over corpora from different time periods). Our analysis yields several interesting findings, e.g., that implicit gender bias in embeddings trained on Arabic news corpora steadily increases over time (between 2007 and 2017). We make the Arabic bias specifications (AraWEAT) publicly available.

📄 PDF Abstract BibTeX arXiv:2011.01575

Code (0)

등록된 구현이 없습니다.

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

LIM-LIG at SemEval-2017 Task1: Enhancing the Semantic Similarity for Arabic Sentences with Vectors Weighting

2017-08-01 · SEMEVAL 2017 8 · El Moatez Billah Nagoudi, J{\'e}r{\'e}my Ferrero, Didier Schwab

This article describes our proposed system named LIM-LIG. This system is designed for SemEval 2017 Task1: Semantic Textual Similarity (Track1). LIM-LIG proposes an innovative enhancement to word embedding-based model dev…

DescriptiveInformation RetrievalMachine TranslationParaphrase Identification+6

Semantic Similarity of Arabic Sentences with Word Embeddings

2017-04-01 · WS 2017 4 · El Moatez Billah Nagoudi, Didier Schwab

Semantic textual similarity is the basis of countless applications and plays an important role in diverse areas, such as information retrieval, plagiarism detection, information extraction and machine translation. This a…

DescriptiveInformation RetrievalMachine TranslationPart-Of-Speech Tagging+8

Are We Consistently Biased? Multidimensional Analysis of Biases in Distributional Word Vectors

2019-04-26 · SEMEVAL 2019 6 · Anne Lauscher, Goran Glavaš

Word embeddings have recently been shown to reflect many of the pronounced societal biases (e.g., gender bias or racial bias). Existing studies are, however, limited in scope and do not investigate the consistency of bia…

Cross-Lingual TransferWord Embeddings

Offline Handwriting Recognition with Multidimensional Recurrent Neural Networks

2008-12-01 · NeurIPS 2008 12 · Alex Graves, Jürgen Schmidhuber

Offline handwriting recognition---the transcription of images of handwritten text---is an interesting task, in that it combines computer vision with sequence learning. In most systems the two elements are handled separat…

Handwriting Recognition

On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena

2025-01-08 · Tarek Naous, Wei Xu

Language Models (LMs) have been shown to exhibit a strong preference towards entities associated with Western culture when operating in non-Western languages. In this paper, we aim to uncover the origins of entity-relate…