paper-with-me

홈 › Papers

Discovering and Categorising Language Biases in Reddit

2020-08-06 · Xavier Ferrer, Tom van Nuenen, Jose M. Such, Natalia Criado

We present a data-driven approach using word embeddings to discover and categorise language biases on the discussion platform Reddit. As spaces for isolated user communities, platforms such as Reddit are increasingly connected to issues of racism, sexism and other forms of discrimination. Hence, there is a need to monitor the language of these groups. One of the most promising AI approaches to trace linguistic biases in large textual datasets involves word embeddings, which transform text into high-dimensional dense vectors and capture semantic relations between words. Yet, previous studies require predefined sets of potential biases to study, e.g., whether gender is more or less associated with particular types of jobs. This makes these approaches unfit to deal with smaller and community-centric datasets such as those on Reddit, which contain smaller vocabularies and slang, as well as biases that may be particular to that community. This paper proposes a data-driven approach to automatically discover language biases encoded in the vocabulary of online discourse communities on Reddit. In our approach, protected attributes are connected to evaluative words found in the data, which are then categorised through a semantic analysis system. We verify the effectiveness of our method by comparing the biases we discover in the Google News dataset with those found in previous literature. We then successfully discover gender bias, religion bias, and ethnic bias in different Reddit communities. We conclude by discussing potential application scenarios and limitations of this data-driven bias discovery method.

📄 PDF Abstract BibTeX arXiv:2008.02754

Code (1)

xfold/LanguageBiasesInReddit 공식 구현

Tasks

Word Embeddings

Similar Papers 제목 키워드 기반

Quantifying Gender Biases Towards Politicians on Reddit

2021-12-22 · Sara Marjanovic, Karolina Stańczak, Isabelle Augenstein

Despite attempts to increase gender parity in politics, global efforts have struggled to ensure equal female representation. This is likely tied to implicit gender biases against women in authority. In this work, we pres…

Bias DetectionGender Bias Detection

IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments

2026-01-10 · Debasmita Panda, Akash Anil, Neelesh Kumar Shukla arxiv

Warning: This paper consists of examples representing regional biases in Indian regions that might be offensive towards a particular region. While social biases corresponding to gender, race, socio-economic conditions, e…

Discovering changes in birthing narratives during COVID-19

2022-04-25 · Daphna Spira, Noreen Mayat, Caitlin Dreisbach, Adam Poliak

We investigate whether, and if so how, birthing narratives written by new parents on Reddit changed during COVID-19. Our results indicate that the presence of family members significantly decreased and themes related to …

Discovering the Functions of Language in Online Forums

2019-11-01 · WS 2019 11 · Youmna Ismaeil, Oana Balalau, Paramita Mirza

In this work, we revisit the functions of language proposed by linguist Roman Jakobson and we highlight their potential in analyzing online forum conversations. We investigate the relationship between functions and other…

Social Meme-ing: Measuring Linguistic Variation in Memes

2023-11-15 · Naitian Zhou, David Jurgens, David Bamman

Much work in the space of NLP has used computational methods to explore sociolinguistic variation in text. In this paper, we argue that memes, as multimodal forms of language comprised of visual templates and text, also …