When do Word Embeddings Accurately Reflect Surveys on our Beliefs About People?
Social biases are encoded in word embeddings. This presents a unique opportunity to study society historically and at scale, and a unique danger when embeddings are used in downstream applications. Here, we investigate the extent to which publicly-available word embeddings accurately reflect beliefs about certain kinds of people as measured via traditional survey methods. We find that biases found in word embeddings do, on average, closely mirror survey data across seventeen dimensions of social meaning. However, we also find that biases in embeddings are much more reflective of survey data for some dimensions of meaning (e.g. gender) than others (e.g. race), and that we can be highly confident that embedding-based measures reflect survey data only for the most salient biases.
Code (1)
Tasks
SurveyWord EmbeddingsSimilar Papers 제목 키워드 기반
SOS: Systematic Offensive Stereotyping Bias in Word Embeddings
Hate speech detection models aim to provide a safe environment for marginalised social groups to express themselves. However, the bias in these models could lead to silencing those groups. In this paper, we introduce the…
Hate Speech DetectionWord EmbeddingsExploring the Value of Personalized Word Embeddings
In this paper, we introduce personalized word embeddings, and examine their value for language modeling. We compare the performance of our proposed prediction model when using personalized versus generic word representat…
Authorship AttributionLanguage ModelingLanguage ModellingWord EmbeddingsSOS: Systematic Offensive Stereotyping Bias in Word Embeddings
Systematic Offensive stereotyping (SOS) in word embeddings could lead to associating marginalised groups with hate speech and profanity, which might lead to blocking and silencing those groups, especially on social media…
BlockingHate Speech DetectionWord EmbeddingsPrediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings
Understanding what knowledge is implicitly encoded in deep learning models is essential for improving the interpretability of AI systems. This paper examines common methods to explain the knowledge encoded in word embedd…
Topology of Word Embeddings: Singularities Reflect Polysemy
The manifold hypothesis suggests that word vectors live on a submanifold within their ambient vector space. We argue that we should, more accurately, expect them to live on a pinched manifold: a singular quotient of a ma…
Word EmbeddingsWord Sense Induction