paper-with-me

Papers

Probing Explicit and Implicit Gender Bias through LLM Conditional Text Generation

2023-11-01 · Xiangjue Dong, Yibo Wang, Philip S. Yu, James Caverlee

Large Language Models (LLMs) can generate biased and toxic responses. Yet most prior work on LLM gender bias evaluation requires predefined gender-related phrases or gender stereotypes, which are challenging to be comprehensively collected and are limited to explicit bias evaluation. In addition, we believe that instances devoid of gender-related language or explicit stereotypes in inputs can still induce gender bias in LLMs. Thus, in this work, we propose a conditional text generation mechanism without the need for predefined gender phrases and stereotypes. This approach employs three types of inputs generated through three distinct strategies to probe LLMs, aiming to show evidence of explicit and implicit gender biases in LLMs. We also utilize explicit and implicit evaluation metrics to evaluate gender bias in LLMs under different strategies. Our experiments demonstrate that an increased model size does not consistently lead to enhanced fairness and all tested LLMs exhibit explicit and/or implicit gender bias, even when explicit gender stereotypes are absent in the inputs.

📄 PDF Abstract BibTeX arXiv:2311.00306

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Text GenerationFairnessText Generation

Similar Papers 제목 키워드 기반

Disclosure and Mitigation of Gender Bias in LLMs

2024-02-17 · Xiangjue Dong, Yibo Wang, Philip S. Yu, James Caverlee

Large Language Models (LLMs) can generate biased responses. Yet previous direct probing techniques contain either gender mentions or predefined gender stereotypes, which are challenging to comprehensively collect. Hence,…

Transcending the "Male Code": Implicit Masculine Biases in NLP Contexts

2023-04-22 · Katie Seaborn, Shruti Chandra, Thibault Fabre

Critical scholarship has elevated the problem of gender bias in data sets used to train virtual assistants (VAs). Most work has focused on explicit biases in language, especially against women, girls, femme-identifying p…

Word Embeddings

Uncovering Implicit Gender Bias in Narratives through Commonsense Inference

2021-09-14 · Findings (EMNLP) 2021 11 · Tenghao Huang, Faeze Brahman, Vered Shwartz, Snigdha Chaturvedi

Pre-trained language models learn socially harmful biases from their training corpora, and may repeat these biases when used for generation. We study gender biases associated with the protagonist in model-generated stori…

AI Will Always Love You: Studying Implicit Biases in Romantic AI Companions

2025-02-27 · Clare Grogan, Jackie Kay, María Pérez-Ortiz

While existing studies have recognised explicit biases in generative models, including occupational gender biases, the nuances of gender stereotypes and expectations of relationships between users and AI companions remai…

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

2026-04-02 · Bhaskara Hanuma Vedula, Darshan Anghan, Ishita Goyal, Ponnurangam Kumaraguru 외 arxiv

Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based pr…