paper-with-me

홈 › Papers

Invisible Influences: Investigating Implicit Intersectional Biases through Persona Engineering in Large Language Models

2026-03-16 · Nandini Arimanda, Achyuth Mukund, Sakthi Balan Muthiah, Rajesh Sharma arxiv

Large Language Models (LLMs) excel at human-like language generation but often embed and amplify implicit, intersectional biases, especially under persona-driven contexts. Existing bias audits rely on static, embedding-based tests (CEAT, I-WEAT, I-SEAT) that quantify absolute association strengths. We show that they have limitations in capturing dynamic shifts when models adopt social roles. We address this gap by introducing the Bias Amplification Differential and Explainability Score (BADx): a novel, scalable metric that measures persona-induced bias amplification and integrates local explainability insights. BADx comprises three components - differential bias scores (BAD, based on CEAT, I-WEAT, I-SEAT),Persona Sensitivity Index (PSI), and Volatility (Standard Deviation), augmented by LIME-based analysis for emphasizing explainability. This study is divided and performed as two different tasks. Task 1 establishes static bias baselines, and Task 2 applies six persona frames (marginalized and structurally advantaged) to measure BADx, PSI, and volatility. This is studied across five state-of-the-art LLMs (GPT-4o, DeepSeek-R1, LLaMA-4, Claude 4.0 Sonnet and Gemma-3n E4B). Results show persona context significantly modulates bias. GPT-4o exhibits high sensitivity and volatility; DeepSeek-R1 suppresses bias but with erratic volatility; LLaMA-4 maintains low volatility and a stable bias profile with limited amplification; Claude 4.0 Sonnet achieves balanced modulation; and Gemma-3n E4B attains the lowest volatility with moderate amplification. BADx performs better than static methods by revealing context-sensitive biases overlooked in static methods. Our unified method offers a systematic way to detect dynamic implicit intersectional bias in five popular LLMs.

📄 PDF Abstract BibTeX arXiv:2604.06213

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

2020-06-06 · Wei Guo, Aylin Caliskan

With the starting point that implicit human biases are reflected in the statistical regularities of language, it is possible to measure biases in English static word embeddings. State-of-the-art neural language models ge…

Bias DetectionSentenceWord Embeddings

How Reasoning Influences Intersectional Biases in Vision Language Models

2025-11-08 · Adit Desai, Sudipta Roy, Mohna Chakraborty arxiv

Vision Language Models (VLMs) are increasingly deployed across downstream tasks, yet their training data often encode social biases that surface in outputs. Unlike humans, who interpret images through contextual and soci…

Explaining the ghosts: Feminist intersectional XAI and cartography as methods to account for invisible labour

2023-05-05 · Goda Klumbyte, Hannah Piehl, Claude Draude

Contemporary automation through AI entails a substantial amount of behind-the-scenes human labour, which is often both invisibilised and underpaid. Since invisible labour, including labelling and maintenance work, is an …

Explainable Artificial Intelligence (XAI)

Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases

2020-10-28 · Ryan Steed, Aylin Caliskan

Recent advances in machine learning leverage massive datasets of unlabeled images from the web to learn general-purpose image representations for tasks from image classification to face recognition. But do unsupervised c…

BIG-bench Machine LearningFace Recognitionimage-classificationImage Classification+1

WordBias: An Interactive Visual Tool for Discovering Intersectional Biases Encoded in Word Embeddings

2021-03-05 · Bhavya Ghai, Md Naimul Hoque, Klaus Mueller

Intersectional bias is a bias caused by an overlap of multiple social factors like gender, sexuality, race, disability, religion, etc. A recent study has shown that word embedding models can be laden with biases against …

Word Embeddings