paper-with-me

Papers

Automated Evaluation of Gender Bias Across 13 Large Multimodal Models

2025-09-08 · Juan Manuel Contreras arxiv

Large multimodal models (LMMs) have revolutionized text-to-image generation, but they risk perpetuating the harmful social biases in their training data. Prior work has identified gender bias in these models, but methodological limitations prevented large-scale, comparable, cross-model analysis. To address this gap, we introduce the Aymara Image Fairness Evaluation, a benchmark for assessing social bias in AI-generated images. We test 13 commercially available LMMs using 75 procedurally-generated, gender-neutral prompts to generate people in stereotypically-male, stereotypically-female, and non-stereotypical professions. We then use a validated LLM-as-a-judge system to score the 965 resulting images for gender representation. Our results reveal (p < .001 for all): 1) LMMs systematically not only reproduce but actually amplify occupational gender stereotypes relative to real-world labor data, generating men in 93.0% of images for male-stereotyped professions but only 22.5% for female-stereotyped professions; 2) Models exhibit a strong default-male bias, generating men in 68.3% of the time for non-stereotyped professions; and 3) The extent of bias varies dramatically across models, with overall male representation ranging from 46.7% to 73.3%. Notably, the top-performing model de-amplified gender stereotypes and approached gender parity, achieving the highest fairness scores. This variation suggests high bias is not an inevitable outcome but a consequence of design choices. Our work provides the most comprehensive cross-model benchmark of gender bias to date and underscores the necessity of standardized, automated evaluation tools for promoting accountability and fairness in AI development.

📄 PDF Abstract BibTeX arXiv:2509.07050

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Deep Generative Views to Mitigate Gender Classification Bias Across Gender-Race Groups

2022-08-17 · Sreeraj Ramachandran, Ajita Rattani

Published studies have suggested the bias of automated face-based gender classification algorithms across gender-race groups. Specifically, unequal accuracy rates were obtained for women and dark-skinned people. To mitig…

ClassificationFacial Attribute ClassificationFairnessGender Classification

Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection

2026-08-04 · Razieh Chalehchaleh, Reza Farahbakhsh, Noel Crespi arxiv

Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context remains underexplored. This study presents the first systematic investigation of gen…

Fake News Detection

Towards Understanding Gender Bias in Relation Extraction

2019-11-09 · ACL 2020 6 · Andrew Gaut, Tony Sun, Shirlyn Tang, Yuxin Huang 외

Recent developments in Neural Relation Extraction (NRE) have made significant strides towards Automated Knowledge Base Construction (AKBC). While much attention has been dedicated towards improvements in accuracy, there …

counterfactualData AugmentationKnowledge Base ConstructionRelation+2

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

2026-03-19 · Shree Harsha Bokkahalli Satish, Maria Teleki, Christoph Minixhofer, Ondrej Klejch 외 arxiv

SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations often miss how such bias manifests in end-to-end speech interactions and h…

Voice Conversion

On Measuring Gender Bias in Translation of Gender-neutral Pronouns

2019-05-28 · WS 2019 8 · Won Ik Cho, Ji Won Kim, Seok Min Kim, Nam Soo Kim

Ethics regarding social bias has recently thrown striking issues in natural language processing. Especially for gender-related topics, the need for a system that reduces the model bias has grown in areas such as image ca…

EthicsImage CaptioningMachine TranslationSentence+1