paper-with-me

홈 › Papers

Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions

2026-05-12 · Jagdish Tripathy, Marcus Buckmann arxiv

Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representations. However, whether these suppressed representations can affect model outputs - and whether such causal potency is symmetric across demographic groups - remains unknown. We investigate the use of open-weight models for mortgage underwriting using matched applications that differ only in racially-associated names and reveal a critical disconnect: models show no output-level bias, yet retain and amplify demographic representations across model layers. Through activation steering and novel cross-layer interventions, we demonstrate that this suppressed information is decision-relevant: when reinjected at critical layers, it produces near-complete decision reversals. Critically, this latent bias is asymmetric - steering interventions affect decisions in one demographic direction, while producing minimal effects in reverse - and susceptible to adversarial prompt engineering and parameter-efficient fine-tuning. These findings demonstrate that behavioural audits focused on outputs are insufficient: fair outputs can mask exploitable internal biases. They also motivate dual-layer testing frameworks combining output evaluation with representational analysis for AI governance in high-stakes decisions.

📄 PDF Abstract BibTeX arXiv:2605.15217

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningPrompt Engineering

Similar Papers 제목 키워드 기반

Fairness Through Causal Awareness: Learning Latent-Variable Models for Biased Data

2018-09-07 · David Madras, Elliot Creager, Toniann Pitassi, Richard Zemel

How do we learn from biased data? Historical datasets often reflect historical prejudices; sensitive or protected attributes may affect the observed treatments and outcomes. Classification algorithms tasked with predicti…

AttributeFairnessGeneral Classification

Fair Clustering: A Causal Perspective

2023-12-14 · Fritz Bayer, Drago Plecko, Niko Beerenwinkel, Jack Kuipers

Clustering algorithms may unintentionally propagate or intensify existing disparities, leading to unfair representations or biased decision-making. Current fair clustering methods rely on notions of fairness that do not …

ClusteringDecision MakingFairness

AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models

2025-09-02 · Snehasis Mukhopadhyay, Aryan Kasat, Shivam Dubey, Rahul Karthikeyan 외 arxiv

Large Language Models (LLMs) can inadvertently reflect societal biases present in their training data, leading to harmful or prejudiced outputs. In the Indian context, our empirical evaluations across a suite of models r…

D-BIAS: A Causality-Based Human-in-the-Loop System for Tackling Algorithmic Bias

2022-08-10 · Bhavya Ghai, Klaus Mueller

With the rise of AI, algorithms have become better at learning underlying patterns from the training data including ingrained social biases based on gender, race, etc. Deployment of such algorithms to domains such as hir…

Fairness

Optimal Experiments for Partial Causal Effect Identification

2026-05-07 · Tobias Maringgele, Jalal Etesami arxiv

Causal queries are often only partially identifiable from observational data, and experiments that could tighten the resulting bounds are typically costly. We study the problem of selecting, prior to observing experiment…