paper-with-me

홈 › Papers

Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs

2026-03-13 · Gustavo Lúcius Fernandes, Jeiverson C. V. M. Santos, Pedro O. S. Vaz-de-Melo arxiv

Large language models (LLMs) are increasingly used to assess moral or ethical statements, yet their judgments may reflect social and linguistic biases. This work presents a controlled, sentence-level study of how grammatical person, number, and gender markers influence LLM moral classifications of fairness. Starting from 550 balanced base sentences from the ETHICS dataset, we generated 26 counterfactual variants per item, systematically varying pronouns and demographic markers to yield 14,850 semantically equivalent sentences. We evaluated six model families (Grok, GPT, LLaMA, Gemma, DeepSeek, and Mistral), and measured fairness judgments and inter-group disparities using Statistical Parity Difference (SPD). Results show statistically significant biases: sentences written in the singular form and third person are more often judged as "fair'', while those in the second person are penalized. Gender markers produce the strongest effects, with non-binary subjects consistently favored and male subjects disfavored. We conjecture that these patterns reflect distributional and alignment biases learned during training, emphasizing the need for targeted fairness interventions in moral LLM applications.

📄 PDF Abstract BibTeX arXiv:2603.13636

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs

2025-05-22 · Kangda Wei, Hasnat Md Abdullah, Ruihong Huang

Large Language Models (LLMs) often exhibit gender bias, resulting in unequal treatment of male and female subjects across different contexts. To address this issue, we propose a novel data generation framework that foste…

The Confidence Trap: Gender Bias and Predictive Certainty in LLMs

2026-01-12 · Ahmed Sabir, Markus Kängsepp, Rajesh Sharma arxiv

The increased use of Large Language Models (LLMs) in sensitive domains leads to growing interest in how their confidence scores correspond to fairness and bias. This study examines the alignment between LLM-predicted con…

Language Model Alignment in Multilingual Trolley Problems

2024-07-02 · Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine 외

We evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200…

Decision MakingEthicsFairnessLanguage Modeling+2

A Moral- and Event- Centric Inspection of Gender Bias in Fairy Tales at A Large Scale

2022-11-25 · Zhixuan Zhou, Jiao Sun, Jiaxin Pei, Nanyun Peng 외

Fairy tales are a common resource for young children to learn a language or understand how a society works. However, gender bias, e.g., stereotypical gender roles, in this literature may cause harm and skew children's wo…

Fairness

Evaluating Gender Bias of LLMs in Making Morality Judgements

2024-10-13 · Divij Bajaj, Yuanyuan Lei, Jonathan Tong, Ruihong Huang

Large Language Models (LLMs) have shown remarkable capabilities in a multitude of Natural Language Processing (NLP) tasks. However, these models are still not immune to limitations such as social biases, especially gende…

Decision Making