paper-with-me

Papers

From Structured Prompts to Open Narratives: Measuring Gender Bias in LLMs Through Open-Ended Storytelling

2025-03-20 · Evan Chen, Run-Jun Zhan, Yan-Bai Lin, Hung-Hsuan Chen

Large Language Models (LLMs) have revolutionized natural language processing, yet concerns persist regarding their tendency to reflect or amplify social biases present in their training data. This study introduces a novel evaluation framework to uncover gender biases in LLMs, focusing on their occupational narratives. Unlike previous methods relying on structured scenarios or carefully crafted prompts, our approach leverages free-form storytelling to reveal biases embedded in the models. Systematic analyses show an overrepresentation of female characters across occupations in six widely used LLMs. Additionally, our findings reveal that LLM-generated occupational gender rankings align more closely with human stereotypes than actual labor statistics. These insights underscore the need for balanced mitigation strategies to ensure fairness while avoiding the reinforcement of new stereotypes.

📄 PDF Abstract BibTeX arXiv:2503.15904

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases

2025-09-04 · Bufan Gao, Elisa Kreiss arxiv

As LLMs are increasingly applied in socially impactful settings, concerns about gender bias have prompted growing efforts both to measure and mitigate such bias. These efforts often rely on evaluation tasks that differ f…

Using Instruction-Tuned Large Language Models to Identify Indicators of Vulnerability in Police Incident Narratives

2024-12-16 · Sam Relins, Daniel Birks, Charlie Lloyd

Objectives: Compare qualitative coding of instruction tuned large language models (IT-LLMs) against human coders in classifying the presence or absence of vulnerability in routinely collected unstructured text that descr…

counterfactualGeneral ClassificationLLM real-life tasksSpecificity

Integrating topic modeling and word embedding to characterize violent deaths

2021-06-28 · Alina Arseniev-Koehler, Susan D. Cochran, Vickie M. Mays, Kai-Wei Chang 외

There is an escalating need for methods to identify latent patterns in text data from many domains. We introduce a new method to identify topics in a corpus and represent documents as topic sequences. Discourse Atom Topi…

ProText: A benchmark dataset for measuring (mis)gendering in long-form texts

2026-03-29 · Hadas Kotek, Margit Bowler, Patrick Sonnenberg, Yu'an Yang arxiv

We introduce ProText, a dataset for measuring gendering and misgendering in stylistically diverse long-form English texts. ProText spans three dimensions: Theme nouns (names, occupations, titles, kinship terms), Theme ca…

Experimental Narratives: A Comparison of Human Crowdsourced Storytelling and AI Storytelling

2023-10-19 · Nina Begus

The paper proposes a framework that combines behavioral and computational experiments employing fictional prompts as a novel tool for investigating cultural artifacts and social biases in storytelling both by humans and …