paper-with-me

홈 › Papers

WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models

2023-06-26 · Virginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan May

We present WinoQueer: a benchmark specifically designed to measure whether large language models (LLMs) encode biases that are harmful to the LGBTQ+ community. The benchmark is community-sourced, via application of a novel method that generates a bias benchmark from a community survey. We apply our benchmark to several popular LLMs and find that off-the-shelf models generally do exhibit considerable anti-queer bias. Finally, we show that LLM bias against a marginalized community can be somewhat mitigated by finetuning on data written about or by members of that community, and that social media text written by community members is more effective than news text written about the community by non-members. Our method for community-in-the-loop benchmark development provides a blueprint for future researchers to develop community-driven, harms-grounded LLM benchmarks for other marginalized communities. Note: This version corrects a bug found in evaluation code after publication. General findings have not changed, but tables 5 and 6 and figure 1 have been corrected.

📄 PDF Abstract BibTeX arXiv:2306.15087

Code (1)

katyfelkner/winoqueer 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Towards WinoQueer: Developing a Benchmark for Anti-Queer Bias in Large Language Models

2022-06-23 · Virginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang, Jonathan May

This paper presents exploratory work on whether and to what extent biases against queer and trans people are encoded in large language models (LLMs) such as BERT. We also propose a method for reducing these biases in dow…

Bias Detection

PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs

2025-07-18 · Maluna Menke, Thilo Hagendorff arxiv

Large Language Models (LLMs) frequently reproduce the gender- and sexual-identity prejudices embedded in their training corpora, leading to outputs that marginalize LGBTQIA+ users. Hence, reducing such biases is of great…

parameter-efficient fine-tuning

Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speech

2025-02-13 · Jonathan Pofcher, Christopher M. Homan, Randall Sell, Ashiqur R. KhudaBukhsh

This paper makes three contributions. First, via a substantial corpus of 1,419,047 comments posted on 3,161 YouTube news videos of major US cable news outlets, we analyze how users engage with LGBTQ+ news content. Our an…

Detecting Harmful Online Conversational Content towards LGBTQIA+ Individuals

2022-06-15 · Jamell Dacon, Harry Shomer, Shaylynn Crum-Dacon, Jiliang Tang

Online discussions, panels, talk page edits, etc., often contain harmful conversational content i.e., hate speech, death threats and offensive language, especially towards certain demographic groups. For example, individ…

Predictive Insights into LGBTQ+ Minority Stress: A Transductive Exploration of Social Media Discourse

2024-11-20 · S. Chapagain, Y. Zhao, T. K. Rohleen, S. M. Hamdi 외

Individuals who identify as sexual and gender minorities, including lesbian, gay, bisexual, transgender, queer, and others (LGBTQ+) are more likely to experience poorer health than their heterosexual and cisgender counte…

Transductive Learning