paper-with-me

홈 › Papers

Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models

2025-05-04 · Paloma Piot, Patricia Martín-Rodilla, Javier Parapar

Commercial Large Language Models (LLMs) have recently incorporated memory features to deliver personalised responses. This memory retains details such as user demographics and individual characteristics, allowing LLMs to adjust their behaviour based on personal information. However, the impact of integrating personalised information into the context has not been thoroughly assessed, leading to questions about its influence on LLM behaviour. Personalisation can be challenging, particularly with sensitive topics. In this paper, we examine various state-of-the-art LLMs to understand their behaviour in different personalisation scenarios, specifically focusing on hate speech. We prompt the models to assume country-specific personas and use different languages for hate speech detection. Our findings reveal that context personalisation significantly influences LLMs' responses in this sensitive area. To mitigate these unwanted biases, we fine-tune the LLMs by penalising inconsistent hate speech classifications made with and without country or language-specific context. The refined models demonstrate improved performance in both personalised contexts and when no context is provided.

📄 PDF Abstract BibTeX arXiv:2505.02252

Code (1)

palomapiot/geographic-bias 공식 구현 pytorch

Tasks

Hate Speech Detection

Similar Papers 제목 키워드 기반

Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF

2024-03-15 · Amey Hengle, Aswini Kumar, Sahajpreet Singh, Anil Bandhakavi 외

Counterspeech, defined as a response to mitigate online hate speech, is increasingly used as a non-censorial solution. Addressing hate speech effectively involves dispelling the stereotypes, prejudices, and biases often …

Sentence

Contextualizing Hate Speech Classifiers with Post-hoc Explanation

2020-05-05 · ACL 2020 6 · Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani 외

Hate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways. Such biases manifest in false positives when these identif…

From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets

2024-04-27 · Manuel Tonneau, Diyi Liu, Samuel Fraiberger, Ralph Schroeder 외

Perceptions of hate can vary greatly across cultural contexts. Hate speech (HS) datasets, however, have traditionally been developed by language. This hides potential cultural biases, as one language may be spoken in dif…

Detecting East Asian Prejudice on Social Media

2020-05-08 · EMNLP (ALW) 2020 11 · Bertie Vidgen, Austin Botelho, David Broniatowski, Ella Guest 외

The outbreak of COVID-19 has transformed societies across the world as governments tackle the health, economic and social costs of the pandemic. It has also raised concerns about the spread of hateful language and prejud…

Identifying and Improving Disability Bias in GPT-Based Resume Screening

2024-01-28 · Kate Glazko, Yusuf Mohammed, Ben Kosa, Venkatesh Potluri 외

As Generative AI rises in adoption, its use has expanded to include domains such as hiring and recruiting. However, without examining the potential of bias, this may negatively impact marginalized populations, including …