paper-with-me

Papers

A Scalable Entity-Based Framework for Auditing Bias in LLMs

2026-01-18 · Akram Elbouanani, Aboubacar Tuo, Adrian Popescu arxiv

Existing approaches to bias evaluation in large language models (LLMs) trade ecological validity for statistical control, relying either on artificial prompts that poorly reflect real-world use or on naturalistic tasks that lack scale and rigor. We introduce a scalable bias-auditing framework that uses named entities as controlled probes to measure systematic disparities in model behavior. Synthetic data enables us to construct diverse, controlled inputs, and we show that it reliably reproduces bias patterns observed in natural text, supporting its use for large-scale analysis. Using this framework, we conduct the largest bias audit to date, comprising 1.9 billion data points across multiple entity types, tasks, languages, models, and prompting strategies. We find consistent patterns: models penalize right-wing politicians and favor left-wing politicians, prefer Western and wealthier countries over the Global South, favor Western companies, and penalize firms in the defense and pharmaceutical sectors. While instruction tuning reduces bias, increasing model scale amplifies it, and prompting in Chinese or Russian does not mitigate Western-aligned preferences. These findings highlight the need for systematic bias auditing before deploying LLMs in high-stakes applications. Our framework is extensible to other domains and tasks, and we make it publicly available to support future work.

📄 PDF Abstract BibTeX arXiv:2601.12374

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When LLMs Imagine People: A Human-Centered Persona Brainstorm Audit for Bias and Fairness in Creative Applications

2026-01-19 · Hongliu Cao, Eoin Thomas, Rodrigo Acuna Agost arxiv

Large Language Models (LLMs) used in creative workflows can reinforce stereotypes and perpetuate inequities, making fairness auditing essential. Existing methods rely on constrained tasks and fixed benchmarks, leaving op…

Bias Detection

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

2026-06-17 · Zilong Zhang, Yi-Ting Hung, Lei Ding, Chi-Kuang Yeh arxiv

Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Me…

Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research

2026-06-10 · Andrés Abeliuk, Cinthia Sanchez Macias, Valentina Alarcón, Álvaro Madariaga 외 arxiv

Research on bias in large language models (LLMs) has predominantly focused on third-person audits, which study how models represent or evaluate demographic groups as external subjects. However, this paradigm overlooks a …

LLMAuditor: A Framework for Auditing Large Language Models Using Human-in-the-Loop

2024-02-14 · Maryam Amirizaniani, Jihan Yao, Adrian Lavergne, Elizabeth Snell Okada 외

As Large Language Models (LLMs) become more pervasive across various users and scenarios, identifying potential issues when using these models becomes essential. Examples of such issues include: bias, inconsistencies, an…

HallucinationTruthfulQA

Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study

2025-10-06 · Ayan Majumdar, Feihao Chen, Jinghui Li, Xiaozhen Wang arxiv

Large-scale web-scraped text corpora used to train general-purpose AI models often contain harmful demographic-targeted social biases, creating a regulatory need for data auditing and developing scalable bias-detection m…

Bias Detection